Start here
Getting into AI safety, step by step
The path I'd hand a friend who wants to work on making AI go well. I'm a few steps down it myself, so this isn't advice from the finish line. It's the map as I understand it, with the resources I trust, in the order I'd do them.
1. Pick a lane#
"AI safety" covers several very different jobs. You don't have to commit forever, but knowing which one you're aiming at decides what you study first.
- Technical alignment research. Interpretability (what's actually happening inside the model), evaluations (measuring dangerous capabilities and propensities), robustness, scalable oversight, and AI control. Heavy on ML, math and experiments.
- AI security and red-teaming. Attacking models and the systems around them: jailbreaks, prompt injection, data poisoning, and, increasingly, agents with real tools and credentials. This is my lane. Security people have a real head start here.
- Governance and policy. Standards, regulation, compute governance, and technical governance work that turns evals into policy. Needs strong writing and judgment more than CUDA.
- Engineering and operations at safety organisations. Research engineering, infrastructure, security for the labs themselves, and the operations work that keeps research teams running.
If you're unsure, do step 2 first. Most people find their lane by noticing which readings they can't stop thinking about.
2. Understand the problem#
Before optimising for a job, understand why people think this matters and where the real disagreements are. BlueDot Impact's courses are the standard on-ramp. They're free, cohort-based and well facilitated. Start with the two-hour Future of AI course, then apply to AGI Strategy or Technical AI Safety.
- BlueDot Impact The Future of AI A free, self-paced two-hour introduction to what AI can do today, where it may be heading, and how people can contribute.
- BlueDot Impact AGI Strategy Cohort-based course on how AI development is going, what good and bad outcomes look like, and what could steer it better.
- BlueDot Impact Technical AI Safety Cohort course for technical people covering alignment and RLHF, interpretability, evaluations and red-teaming, AI control, and scalable oversight.
- BlueDot Impact Frontier AI Governance Roughly 40-hour cohort course on frontier AI governance: weighing evidence, comparing policy strategies, and finding where you can contribute.
- 80,000 Hours The risks of advanced AI (problem profile) Long-form explainer of why advanced AI could pose catastrophic risks and why 80,000 Hours considers the problem neglected.
- 80,000 Hours AI safety technical research (career review) Career review covering what technical AI safety research involves, the skills employers look for, and common routes into the field.
- French Center for AI Safety (CeSIA) AI Safety Atlas Free online textbook covering AI capabilities, risks, strategies, governance, evaluations, specification gaming, goal misgeneralization, and scalable oversight.
- Robert Miles Robert Miles AI Safety Accessible YouTube explainers of core AI alignment and safety concepts, aimed at a general audience.
- Center for AI Safety Introduction to AI Safety, Ethics, and Society Textbook by Dan Hendrycks, free to read online, on how AI systems work, societal-scale risks, and approaches to managing them.
- AISafety.com AISafety.com Directory hub for the AI safety field: courses, training programs, events, communities, jobs, funding sources, and free 1-to-1 advisors.
- AISafety.com AISafety.com field map Visual map of hundreds of organizations, programs, and resources across AI safety research, training, funding, governance, and media.
3. Build the foundations#
You need less math than you fear and more than you'd like. The core is linear algebra, multivariable calculus (really just the chain rule and gradients), probability and statistics, and enough optimization to understand gradient descent. Then build neural networks from scratch until the magic is gone.
- 3Blue1Brown Essence of Linear Algebra Short animated video series that builds geometric intuition for vectors, matrices, linear transformations, determinants, and eigenvectors.
- 3Blue1Brown Neural networks Animated series explaining how neural networks learn, including gradient descent and backpropagation, with later chapters on transformers and attention.
- MIT OpenCourseWare MIT 18.06 Linear Algebra (Gilbert Strang) Gilbert Strang's classic MIT linear algebra course with full lecture videos, free on OpenCourseWare.
- Khan Academy Multivariable calculus Free course on partial derivatives, gradients, the Jacobian and Hessian, multiple integrals, and Green's, Stokes', and divergence theorems.
- Harvard University (Joe Blitzstein) Statistics 110: Probability Joe Blitzstein's Harvard probability course: lecture videos, handouts, and a free online edition of the companion textbook.
- Andrej Karpathy Neural Networks: Zero to Hero Free video course that builds neural networks from scratch in Python, from backpropagation (micrograd) through character-level models (makemore) to GPT and tokenizers.
- Zhang, Lipton, Li, and Smola Dive into Deep Learning Free interactive deep learning textbook with runnable code in PyTorch, JAX, TensorFlow, and MXNet, used in courses at hundreds of universities.
- Deisenroth, Faisal, and Ong (Cambridge University Press) Mathematics for Machine Learning Free textbook covering the linear algebra, calculus, probability, and optimization behind machine learning, then applying them to core ML methods.
- fast.ai Practical Deep Learning for Coders Free, code-first deep learning course for people who can program, covering vision, NLP, tabular data, and deployment.
My rule: learn each piece of math the week before you need it, not in the abstract. Linear algebra is the week before attention, statistics is the week before evaluation, and optimization is the week before adversarial attacks. My roadmap is built this way if you want an example.
4. Get hands-on with safety engineering#
This is where it turns from "interested in AI safety" into "can do AI safety work". Everything here is free.
- ARENA (Alignment Research Engineer Accelerator) ARENA curriculum Free self-study curriculum of coding exercises: Fundamentals, Transformer Interpretability, Reinforcement Learning, LLM Evaluations, and a newer Alignment Science chapter.
- Neel Nanda How To Become A Mechanistic Interpretability Researcher Opinionated guide to learning mechanistic interpretability: learn the basics quickly, practice with short mini-projects, then work up to full research projects.
- Anthropic Transformer Circuits Thread Anthropic's interpretability research publication, from 'A Mathematical Framework for Transformer Circuits' to recent work on features and circuits in Claude.
- TransformerLensOrg (created by Neel Nanda) TransformerLens Open-source Python library for mechanistic interpretability that loads open language models and exposes their internal activations for inspection and editing.
- NDIF (Northeastern University) nnsight Python library for reading and intervening on the internals of deep learning models; the same code can run on large remote models via NDIF.
- UK AI Security Institute Inspect Open-source framework for building and running LLM evaluations, including agentic tasks, with a catalog of over 200 ready-made benchmarks.
- Neuronpedia Neuronpedia Open-source interpretability platform to browse sparse-autoencoder features, trace circuits with attribution graphs, and steer model activations in the browser.
ARENA is the single highest-leverage resource on this page if you're going technical. It takes you from PyTorch fundamentals to transformer interpretability, reinforcement learning and evals, with exercises that make you actually implement things.
5. The security lane#
If you're coming from security, as I did, lean into it. The field needs people who think like attackers, and most ML people don't. Learn how the attacks work, then learn how they're measured, because evaluation is where most of the hard problems hide. If you're a security professional, look at the free, fully funded AI Security Bootcamp (listed under programs in step 8). It's built for exactly this move.
Learn the attack surface#
- 80,000 Hours AI security (career review) Career review on why security skills matter for AI: protecting AI systems from attackers, containing misaligned AI, and supporting verification and governance.
- Hack The Box Academy (with Google) AI Red Teamer job-role path Hands-on 12-module path, built with Google and aligned to its SAIF framework, covering prompt injection, adversarial ML, data and privacy attacks, and defense.
- PortSwigger Web LLM attacks (Web Security Academy) Free lessons and hands-on labs on attacking LLM-integrated web apps, such as indirect prompt injection and insecure output handling.
- OWASP GenAI Security Project OWASP Top 10 for LLM Applications 2026 Community-ranked list of the ten most critical security risks for LLM applications, from prompt injection to improper output handling, with mitigations.
- OWASP GenAI Security Project OWASP Top 10 for Agentic Applications (2026) OWASP's list of the top ten security risks for autonomous AI agents that plan, use tools, keep memory, and act on users' behalf.
- MITRE MITRE ATLAS ATT&CK-style knowledge base of adversary tactics, techniques, and real-world case studies for attacks on AI-enabled systems.
- Google Secure AI Framework (SAIF) Google's practitioner guide to AI security, with a risk map, controls for common AI risks, agent guidance, and a self-assessment.
- NIST NIST AI 100-2 E2025: Adversarial Machine Learning NIST's taxonomy and shared vocabulary for attacks on predictive and generative AI, such as evasion, poisoning, privacy, and prompt injection, plus mitigations.
- NIST NIST AI Risk Management Framework (AI 100-1) NIST's voluntary framework for identifying and managing AI risks across the AI lifecycle; a revision is in progress.
- NIST NIST AI 600-1: Generative AI Profile Companion profile to the AI RMF describing risks specific to generative AI and suggested actions for managing them.
Practice legally#
- Gray Swan AI Gray Swan Arena Free AI red-teaming arena with ongoing practice challenges and sponsored competitions that pay cash prizes for successful jailbreaks and injections.
- HackAPrompt / Learn Prompting HackAPrompt Free prompt-hacking competition platform with beginner tutorial tracks and themed challenge tracks for jailbreaking and indirect prompt injection.
- Lakera (a Check Point company) Gandalf / Agent Breaker Browser game for learning prompt injection; the Gandalf address now leads to Agent Breaker, which targets realistic AI agent apps.
Tools and benchmarks#
- NVIDIA garak Open-source LLM vulnerability scanner that probes models for jailbreaks, prompt injection, data leakage, toxicity, and hallucination.
- Microsoft PyRIT Microsoft's open-source Python framework that helps security professionals and engineers proactively find risks in generative AI systems.
- Promptfoo (part of OpenAI) promptfoo Open-source CLI and library for testing prompts, agents, and RAG apps, including automated red-teaming and vulnerability scanning.
- Center for AI Safety HarmBench Standardized framework for evaluating automated red-teaming attacks and how robustly LLMs refuse harmful requests.
- JailbreakBench team JailbreakBench Open benchmark and leaderboard for jailbreak attacks and defenses, with a 100-behavior misuse dataset and a standardized evaluation library.
- Souly et al. StrongREJECT Jailbreak evaluation benchmark and automated grader designed to avoid overstating attack success; ships dozens of baseline jailbreak implementations.
- ETH Zurich SPY Lab AgentDojo Benchmark environment for testing prompt injection attacks and defenses against tool-using LLM agents across realistic tasks.
Bug bounties that cover AI#
Read each program's scope carefully. Several explicitly exclude plain jailbreaks and pay for real security impact, like data exfiltration or rogue agent actions.
- Anthropic Anthropic bug bounty programs Public HackerOne program for security vulnerabilities in Anthropic systems; a separate invite-only program pays for universal jailbreaks of Claude's safeguards.
- OpenAI OpenAI Safety Bug Bounty Bugcrowd program rewarding reproducible safety and abuse issues in OpenAI products, such as third-party prompt injection and harmful agent actions.
- Google AI Vulnerability Reward Program Google's AI bug bounty for security-impacting issues like rogue actions and data exfiltration; plain jailbreaks and content issues are out of scope.
- 0DIN (Mozilla) 0DIN Mozilla-backed GenAI bug bounty that accepts reports of model vulnerabilities such as prompt injection and jailbreaks, alongside commercial AI security products.
A word on ethics, because it matters here: attack models you own or are explicitly authorised to test, respect each program's rules, and disclose responsibly. Don't publish working bypasses for deployed systems. The goal is fewer successful attacks in the world, not more.
6. Read papers like it's your job#
Pick one paper a week and read it properly: abstract and figures first, then the method, then the part where you try to reproduce the key number in your head. Write a short summary in your own words. The summaries compound. These are the places I find what's worth reading:
Forums and preprints#
- AI Alignment Forum AI Alignment Forum Forum for technical AI alignment research write-ups and discussion, free to read.
- LessWrong LessWrong Community blog where much AI safety and alignment writing appears first, alongside broader posts on reasoning and rationality.
- alphaXiv alphaXiv Research discovery site for arXiv papers, with trending papers, researcher pages, and AI-assisted literature review.
- arXiv arXiv: Machine Learning (cs.LG) Daily list of new machine learning preprints on arXiv.
- arXiv arXiv: Computation and Language (cs.CL) Daily list of new natural language processing and language model preprints on arXiv.
- arXiv arXiv: Cryptography and Security (cs.CR) Daily list of new security preprints on arXiv, including many papers on attacking and defending AI systems.
Labs and institutes#
- Anthropic Alignment Science Blog Research notes and early findings from Anthropic's Alignment Science team on topics like reward hacking, alignment faking, and auditing.
- Anthropic Anthropic research Anthropic's research hub spanning alignment, interpretability, societal impacts, economics, and its Frontier Red Team.
- OpenAI OpenAI Alignment Research Blog Informal research updates from OpenAI's alignment team on evaluations, monitoring, misalignment detection, and chain-of-thought.
- Google DeepMind AGI Safety & Alignment Research at GDM Blog from Google DeepMind's AGI Safety and Alignment team sharing research ideas and periodic summaries of the team's work.
- METR (Model Evaluation & Threat Research) METR Research nonprofit that measures whether and when AI systems might threaten catastrophic harm; publishes evaluations, research, and technical notes.
- Apollo Research Apollo Research Research organization studying AI scheming, where models pursue hidden goals while appearing aligned, and methods to detect and prevent it.
- Redwood Research Redwood Research blog Blog of Redwood Research, a nonprofit focused on AI control and other techniques to mitigate catastrophic AI risks.
- UK Department for Science, Innovation and Technology UK AI Security Institute (AISI) UK government research organization that tests frontier AI for security risks; publishes research, evaluations, and a blog.
- NIST (US Department of Commerce) Center for AI Standards and Innovation (CAISI) NIST center that evaluates frontier AI capabilities with national security implications and develops voluntary AI standards and guidance.
Newsletters and security blogs#
- Center for AI Safety AI Safety Newsletter Newsletter from the Center for AI Safety summarizing recent AI safety news and policy developments.
- Jack Clark Import AI Jack Clark's weekly newsletter analyzing notable AI research and its implications.
- Simon Willison Simon Willison: prompt injection Long-running archive of posts tracking prompt injection attacks, defenses, and incidents in LLM tools and agents.
- Johann Rehberger Embrace The Red Security researcher's blog documenting real-world prompt injection and data exfiltration exploits against AI assistants and coding agents.
7. Build something and publish it#
Nothing on this page counts as much as a public project that shows how you think. Good first projects:
- Replicate a paper's key result on a small model, and write up what did and didn't reproduce.
- Measure something carefully. Grader disagreement, refusal rates under different framings, or how an attack's success changes with its budget. Report error bars.
- Build a small tool people in the field would actually use, then document it properly.
Write it up where people will see it: a blog post, LessWrong or the Alignment Forum, a GitHub README with real results. A small, honest, well-measured project beats an ambitious unfinished one every time.
8. Apply to programs#
Structured programs give you mentorship, a cohort, and often a stipend. They're competitive, and a public project from step 7 is often what gets you in. Check each site for current dates and eligibility.
- MATS MATS Mentored AI safety research fellowship in Berkeley and London, with stipend, housing, compute, and possible funded extensions.
- SPAR (Research Program for AI Risks) SPAR Part-time, remote, three-month research program pairing mentees with AI safety and policy mentors; unpaid, but project costs are covered.
- AI Safety Camp AI Safety Camp Three-month online, part-time program where teams work on preselected AI safety projects; generally no stipends.
- London AI Safety Research (LASR) Labs LASR Labs Full-time, in-person 13-week program in London where small teams write an AI safety research paper under supervision; stipend provided.
- Pivotal Research Pivotal AI Safety Fellowship Fully funded, in-person AI safety research fellowship in London with mentorship, stipend, and travel and housing support.
- ERA (Cambridge, UK) Cambridge ERA:AI Fellowship Paid in-person research fellowship in Cambridge, UK, with technical AI safety, AI governance, and technical AI governance streams.
- Anthropic Anthropic Fellows Program Four-month funded fellowship for empirical AI safety research with Anthropic mentors; areas include interpretability, AI control, and AI security.
- Constellation Institute Astra Fellowship Fully funded five-month in-person fellowship in Berkeley pairing fellows with senior mentors, with empirical and strategy-and-governance streams.
- Cambridge Boston Alignment Initiative CBAI AI Safety Research Fellowship Fully funded, in-person AI safety research fellowship in Cambridge, Massachusetts, with mentorship, compute, and a stipend.
- ARENA ARENA in-person bootcamp Four-to-five-week in-person ML engineering bootcamp in London built on the ARENA curriculum; travel, accommodation, and meals are covered.
- OpenAI OpenAI Safety Fellowship Pilot fellowship for external researchers to do independent AI safety research with OpenAI mentorship, a stipend, and compute; Berkeley-based or remote.
- AISB (fiscally sponsored by BlueDot Impact) AI Security Bootcamp (AISB) Free, fully funded week-long intensive teaching security professionals adversarial ML, LLM security, AI infrastructure security, and threat modeling.
- BlueDot Impact Technical AI Safety Project sprint Free, part-time project sprint (about 30 hours) where you build a technical AI safety project with weekly expert check-ins.
- Apart Research Apart Research sprints Remote AI safety research hackathons open to anyone, plus fellowships that help turn promising sprint projects into publications.
- AISafety.com AISafety.com training directory Continuously updated directory of AI safety fellowships, bootcamps, and courses, filterable by focus, entry level, stipend, and location.
9. Find your people and your funding#
This field is small and unusually generous with its time. Join a local group, go to events, and share your work where people can respond to it.
- University of Wisconsin-Madison Wisconsin AI Safety Initiative (WAISI) University of Wisconsin-Madison group running AI safety fellowships, reading groups, research, and public speaker events.
- AISafety.com AISafety.com community directory (incl. AI Alignment Slack) Directory of hundreds of AI safety groups; includes the AI Alignment Slack, a large online community for introductions and study partners.
- Centre for Effective Altruism Effective Altruism Forum Discussion forum for the effective altruism community, with an active AI safety topic covering careers, research, and funding.
- Centre for Effective Altruism EA Global and EAGx conferences Application-based conferences for networking and one-on-one career conversations; discounted or free tickets and travel support are available on request.
- AISafety.com AISafety.com events Calendar of AI safety conferences, talks, workshops, and meetups, online and in person, including sessions around major ML conferences.
- AI Village (DEF CON community) AI Village DEF CON community for AI security that runs generative AI red-team events, workshops, and labs, with an active Discord.
If you need runway to make the switch, some funders specifically support career transitions into AI safety. The landscape changed in 2026: the Long-Term Future Fund closed, and Coefficient Giving (formerly Open Philanthropy) handed its career-transition grants to BlueDot. These are the current routes:
- BlueDot Impact Career Transition Grants Funding for a defined full-time period to transition into AI safety or biosecurity work, plus introductions and peer support.
- BlueDot Impact Rapid Grants Small, fast grants (up to $20k) for next steps in AI safety or biosecurity, such as living costs, compute, events, or travel.
- EA Funds Transformative AI Fund Rolling-application fund for early-stage AI safety and governance work, open to individuals and new projects; replaced the Long-Term Future Fund.
- Manifund Manifund Open platform where people post charitable project proposals for public fundraising or for funding from expert regrantors.
Your first 30 days#
If you want something concrete, here's what I'd do starting today:
- Week 1. Take BlueDot's free two-hour Future of AI course and apply to their next cohort, then watch 3Blue1Brown's linear algebra series.
- Week 2. Build micrograd from scratch with Karpathy. Start one paper a week.
- Week 3. Start ARENA Chapter 0, or the Hack The Box AI red-teaming path if you're in the security lane.
- Week 4. Pick one small measurement project and publish the result, however small.
Then do it again, harder. And if you want company, follow along. I'm doing the same thing in public.
Every link on this page was checked on September 27, 2026. Programs change fast, so check dates and eligibility on each site. If something here is out of date, tell me.