FAR.AI · Remote, Global · about 11 hours ago
FAR.AI http://FAR.AI is seeking a Jailbreaking Lead (Red Team) whose personal mission and obsession is to jailbreak the world's leading frontier AI models. You will sit at the tip of the spear of one of the world's leading AI red-teams, with a single, critical focus: find the universal jailbreaks that no one else can find, in the models used by hundreds of millions of people, and make sure they get fixed.
This is primarily a senior IC role with some management responsibilities, ideally for candidates who want to build and lead a jailbreaking team over time. An IC-only track is also available. Either way, you will spend the majority of your time hands-on, building attacks and breaking frontier models, and setting the technical bar for what a world-class jailbreak looks like.
About FAR.AI
FAR.AI is a non-profit AI research institute dedicated to ensuring advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.
Since our founding in July 2022, we've grown quickly to 45+ staff https://www.far.ai/about/team, producing over 40 influential academic papers https://scholar.google.com/citations?user=FVJ24k8AAAAJ, and establishing leading AI Safety events https://far.ai/events/. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML, and ICLR, and features in the Financial Times https://www.ft.com/content/175e5314-a7f7-4741-a786-273219f433a1, Nature News https://www.nature.com/articles/d41586-024-02218-7 and MIT Technology Review https://www.technologyreview.com/2020/02/28/905615/reinforcement-learning-adversarial-attack-gaming-ai-deepmind-alphazero-selfdriving-cars/. Additionally, we help steer and grow the AI safety field through developing https://arxiv.org/abs/2405.06624research https://arxiv.org/abs/2506.20702roadmaps https://www.researchgate.net/publication/396910034_Open_Technical_Problems_in_Open-Weight_AI_Model_Risk_Management with renowned researchers such as Yoshua Bengio; running FAR.Labs https://www.far.ai/programs/far-labs, an AI safety-focused co-working space in Berkeley housing 40 members; and supporting the community through targeted grants https://www.far.ai/programs/grantmaking to technical researchers.
ABOUT RED-TEAMING AT FAR.AI
FAR.AI’s red team is building toward a simple outcome: materially raising the bar for safety and security of the most widely deployed and capable AI systems in the world. We intend to be the tip of the spear in AI safety: the team that consistently finds the failures others miss, resulting in real mitigations, and setting the standard that labs and governments converge on. We also leverage our in-depth understanding of weaknesses in frontier models to advise frontier developers on mitigations, to guide our own research and grant-making for improving model security, and to inform the public of key AI risks.
We are already one of the leading independent red-teaming organizations. Our work has helped most Western frontier model developers improve safeguards through pre- and post-deployment testing (e.g., we have directly influenced safeguards at major frontier developers like OpenAI and Anthropic), and we are increasingly embedded in high-leverage government efforts (e.g., leading a consortium building CBRN evaluations for the European Commission/EU AI Office, and collaborating with the UK AI Security Institute).
"FAR.AI's pre-deployment testing of GPT-5 series models identified failure modes and mitigations, improving the security of our model releases." – Senior Technical Program Manager, OpenAI
“FAR.AI have been a trusted and thoughtful collaborator for us, and they have progressed the state of frontier red-teaming through research like STACK. We expect this to be a high impact role and are excited to explore collaborations with the successful candidate.” – Xander Davies, Technical Lead, Red Team at UK AISI
You will be the senior technical owner of our jailbreaking practice reporting to Kellin Pelrine https://www.far.ai/about/people/kellin-pelrine with a dotted line to Edward Yee https://www.far.ai/about/people/edward-yee. In 2026, we are scaling from a strong team with standout wins into a new level of impact for any AI red team globally:
Jailbreaking is the core technical engine of the red team. As Jailbreaking Lead, you own that engine. You are the person who personally breaks the hardest targets, sets the bar the rest of the team pushes toward, and makes sure we keep discovering the highest severity, universal vulnerabilities – the most important vulnerabilities to fix – in the most heavily defended frontier models on the planet, faster than anyone else.
We expect you to spend at least 50-70% of your time hands-on across 2026: breaking models, chaining novel attack classes through defense-in-depth stacks, helping to invent new techniques when existing ones fail, and setting the standard for what constitutes a significant vulnerability and a credible mitigation. The remaining time will go to managing/mentoring ICs, helping to shape the jailbreaking research agenda with Kellin, and making sure our findings land with frontier labs, governments, and the broader field. The rest of the red team will empower your work, whether through direct collaboration and support, novel research and red-teaming infrastructure, or toolkits and agent build-outs.
This is a senior IC role by default, intended to attract a world-class jailbreaker whose personal mission is to find critical jailbreaks in the most heavily defended domains of the leading frontier AI models, and who has a track record of repeatedly doing so. We are open to a management track for candidates who want to hire and lead a jailbreaking team over time. We will not water down the IC bar to support the management track: both versions of this role require you to be, or be on a clear trajectory to being, one of the best jailbreakers in the world.
If you are earlier in your career but have a standout jailbreaking track record, we still encourage you to apply as we are also hiring for less senior positions on the team. If you are missing the hands-on jailbreaking depth but have adjacent strengths, we encourage you to consider our other open roles on the red team.
If based in the USA or Singapore, you will be an employee of FAR.AI (501(c)(3) research non-profit / non-profit CLG). Outside the USA or Singapore, you will be employed via an EOR organization on behalf of FAR.AI.
We know these roles are rare and the skill combination is unusual. If you're uncertain whether your background fits but are excited by the mission and challenges, we encourage you to apply – we're looking for excellence and potential, not a perfect resume match.
Headquarters
Remote
Work Location
remote
Job Category
Not specified
Application Deadline
Not specified
Job Type
full-time
Experience Level
lead
Application Method
Apply via Website
Salary
170k - 250k USD
No related jobs found