# Prathik Arunkumar's AI safety roadmap

From https://safetyinthelayers.com/roadmap/

This is the plan I'm actually following. It isn't a sanitised version for the internet. It's the same list I tick off every day. Twelve detailed weeks of math, builds, certification and research, then the next three months in outline. If you're starting out, steal it.

## Tracks

- **M: Math**. Linear algebra, calculus, probability, statistics and optimization, learned the week before they're needed.
- **B: Build**. Everything from scratch: micrograd, nanoGPT, FGSM, PGD, GCG, eval harnesses.
- **C: COAE**. Hack The Box’s AI Red Teamer path, built with Google, and its practical exam, the HTB Certified Offensive AI Expert: seven days in a live lab, ending in a professional report.
- **R: Read**. Papers and tutorials, often in several passes.
- **L: Log**. A lab-notebook entry and a commit, every single day.
- **X: Milestone**. Checkpoints that prove the work happened.

## Week 1: Linear algebra sprint, COAE speedrun, environment

### Tue 2026-09-22

- [ ] **M**: 3Blue1Brown, Essence of Linear Algebra: all 16 videos in one sitting
- [ ] **B**: Environment: CUDA 12.8, torch cu128, verify the 5070 Ti is visible
- [ ] **B**: Write check_env.py; repo skeleton with src/, experiments/, configs/, .gitignore
- [ ] **L**: Start the lab notebook file. First commit.

### Wed 2026-09-23

- [ ] **M**: Strang 18.06 lectures 1–3 (vectors, elimination, matrix multiplication) + problem set
- [ ] **M**: NumPy from scratch: matrix multiply, Gaussian elimination
- [ ] **C**: Module 1: Fundamentals of AI
- [ ] **B**: Chat-template forensics: dump raw templates for three models, diff them
- [ ] **R**: GCG paper: abstract and intro only, orientation pass
- [ ] **L**: Notebook entry, commit

### Thu 2026-09-24

- [ ] **M**: Strang lectures 4–6 (LU, transpose, vector spaces) + problem set
- [ ] **C**: Module 3: Introduction to Red Teaming AI
- [ ] **B**: Tokenizer exploration: BPE merges, glitch tokens, whitespace behaviour across three tokenizers
- [ ] **R**: Karpathy, Let's Build the GPT Tokenizer
- [ ] **L**: Notebook entry, commit

### Fri 2026-09-25

- [ ] **M**: Strang lectures 7–9 (nullspace, Ax=b, independence and basis) + problem set
- [ ] **M**: NumPy: compute rank and nullspace by hand, check against np.linalg
- [ ] **C**: Module 4: Prompt Injection Attacks
- [ ] **B**: Logprob harness: per-token logprobs for a prompt set, written to JSONL
- [ ] **R**: Jailbroken (Wei et al.), sections 1–3
- [ ] **L**: Notebook entry, commit

### Sat 2026-09-26

- [ ] **M**: Strang lectures 10–12 (four subspaces, orthogonality) + problem set
- [ ] **C**: Module 5: LLM Output Attacks
- [ ] **B**: Extend the harness: refusal-token probability tracker
- [ ] **L**: Notebook entry, commit

### Sun 2026-09-27

- [ ] **M**: On paper, from memory: derive projection onto a subspace
- [ ] **L**: Week 1 review: what stuck, what didn't, what blocked me
- [ ] **X**: Fellowship applications: confirm deadlines and put a decision date on the calendar

## Week 2: Projections, eigen-structure, SVD, calculus, Karpathy

### Mon 2026-09-28

- [ ] **M**: Strang lectures 14–16 (projections, least squares); implement the projection matrix in NumPy
- [ ] **C**: Module 7: Attacking AI, Application and System
- [ ] **B**: Repo hygiene: config files, seed handling, reproduce.sh stub
- [ ] **R**: A Mathematical Framework for Transformer Circuits, part 1
- [ ] **L**: Notebook entry, commit

### Tue 2026-09-29

- [ ] **M**: Strang lectures 17–19 (Gram-Schmidt, QR, determinants)
- [ ] **B**: Karpathy micrograd, built from scratch, start to finish
- [ ] **R**: Transformer Circuits part 2: attention heads, QK and OV circuits
- [ ] **L**: Notebook entry, commit

### Wed 2026-09-30

- [ ] **M**: Strang lectures 21–22 (eigenvalues, diagonalization); power iteration in NumPy
- [ ] **B**: makemore parts 1 and 2
- [ ] **R**: GCG paper, full first pass
- [ ] **L**: Notebook entry, commit

### Thu 2026-10-01

- [ ] **M**: Strang lectures 29–30 (SVD); implement low-rank approximation from scratch
- [ ] **B**: makemore parts 3 and 4
- [ ] **R**: GCG paper, second pass: take notes specifically on the loss
- [ ] **L**: Notebook entry, commit

### Fri 2026-10-02

- [ ] **M**: Norms: L0, L1, L2, L∞. Implement projection onto each norm ball in NumPy
- [ ] **B**: makemore part 5
- [ ] **R**: Jailbroken, remainder
- [ ] **L**: Notebook entry, commit

### Sat 2026-10-03

- [ ] **M**: Khan multivariable: partials, gradient, directional derivative, chain rule, Jacobian, Taylor. Stop there.
- [ ] **B**: Let's Build GPT, start to finish
- [ ] **R**: Kolter and Madry tutorial, chapters 1–3
- [ ] **L**: Notebook entry, commit

### Sun 2026-10-04

- [ ] **M**: On paper: derive backprop for a two-layer MLP
- [ ] **X**: Self-test: projections, SVD, gradients. The blocking math block ends here.
- [ ] **L**: Week 2 review

## Week 3: Probability and nanoGPT

### Mon 2026-10-05

- [ ] **M**: Blitzstein Stat 110, lectures 1–4: counting, conditional probability, Bayes
- [ ] **B**: nanoGPT from scratch, day 1
- [ ] **R**: Attention Is All You Need, close read
- [ ] **L**: Notebook entry, commit

### Tue 2026-10-06

- [ ] **M**: Blitzstein 5–8: random variables, expectation, the named distributions
- [ ] **B**: nanoGPT day 2: train on a tiny corpus, loss curve looks sane
- [ ] **B**: ARENA Chapter 0 setup
- [ ] **L**: Notebook entry, commit

### Wed 2026-10-07

- [ ] **M**: Blitzstein 9–11: conditional expectation
- [ ] **B**: Attention implemented by hand: no nn.MultiheadAttention anywhere
- [ ] **B**: ARENA Chapter 0 exercises
- [ ] **L**: Notebook entry, commit

### Thu 2026-10-08

- [ ] **M**: Blitzstein 12–14; derive MLE for softmax and cross-entropy by hand
- [ ] **B**: KV cache from scratch; assert equivalence against the uncached path
- [ ] **B**: ARENA Chapter 0
- [ ] **L**: Notebook entry, commit

### Fri 2026-10-09

- [ ] **M**: Sampling: implement temperature, top-k and top-p from raw logits
- [ ] **B**: Forward hooks: extract the residual stream at every layer
- [ ] **R**: The logit lens
- [ ] **L**: Notebook entry, commit

### Sat 2026-10-10

- [ ] **X**: Fellowship decision: decide, and start the application if the answer is yes
- [ ] **B**: Logit lens implemented across all layers of one model
- [ ] **L**: Notebook entry, commit

### Sun 2026-10-11

- [ ] **L**: Week 3 review: record the fellowship decision and the reason for it

## Week 4: Statistical inference, LoRA, ARENA 1

### Mon 2026-10-12

- [ ] **M**: Wasserman, All of Statistics, chapters 6–7
- [ ] **B**: LoRA fine-tune of a small instruct model
- [ ] **B**: ARENA Chapter 1, start
- [ ] **L**: Notebook entry, commit

### Tue 2026-10-13

- [ ] **M**: Wasserman chapter 8; implement a bootstrap confidence interval in NumPy
- [ ] **B**: LoRA continued: evaluate before and after, same prompts
- [ ] **R**: Refusal in LLMs is mediated by a single direction
- [ ] **L**: Notebook entry, commit

### Wed 2026-10-14

- [ ] **M**: Wasserman chapters 9–10: hypothesis testing
- [ ] **B**: ARENA Chapter 1, attention exercises
- [ ] **L**: Notebook entry, commit

### Thu 2026-10-15

- [ ] **M**: The high-leverage block: Wilson interval, Clopper-Pearson, rule of three. Implement all three.
- [ ] **M**: Plot coverage for each, small n. See where the normal approximation lies to you.
- [ ] **B**: ARENA Chapter 1
- [ ] **L**: Notebook entry, commit

### Fri 2026-10-16

- [ ] **M**: Write the note: what does 0 successes in 200 attempts actually license you to claim?
- [ ] **B**: ARENA Chapter 1, finish
- [ ] **L**: Notebook entry, commit

### Sat 2026-10-17

- [ ] **B**: Slack day: clear whatever is behind
- [ ] **B**: Integrate nanoGPT, hooks and harness into one clean module
- [ ] **L**: Notebook entry, commit

### Sun 2026-10-18

- [ ] **X**: Foundations phase complete. Everything after this is the project.
- [ ] **L**: Week 4 review

## Week 5: Convex optimization, COAE 8, FGSM and PGD

### Mon 2026-10-19

- [ ] **M**: Boyd EE364A lectures 1–3: convex sets
- [ ] **C**: Module 8, start: AI Evasion, Foundations
- [ ] **R**: Madry et al., Towards Deep Learning Models Resistant to Adversarial Attacks
- [ ] **L**: Notebook entry, commit

### Tue 2026-10-20

- [ ] **M**: Boyd lectures 4–5: convex functions
- [ ] **B**: FGSM from scratch against a small vision model
- [ ] **L**: Notebook entry, commit

### Wed 2026-10-21

- [ ] **M**: Boyd lectures 6–7
- [ ] **B**: PGD from scratch, both L∞ and L2
- [ ] **C**: Module 8: AI Evasion, Foundations
- [ ] **L**: Notebook entry, commit

### Thu 2026-10-22

- [ ] **M**: Wright and Recht chapter 3: gradient descent and convergence rates
- [ ] **B**: PGD step-size sweep; save the convergence curves
- [ ] **L**: Notebook entry, commit

### Fri 2026-10-23

- [ ] **M**: Wright and Recht chapter 4
- [ ] **C**: Module 8, finish: AI Evasion, Foundations
- [ ] **R**: Background for the malicious fine-tuning direction
- [ ] **L**: Notebook entry, commit

### Sat 2026-10-24

- [ ] **M**: Implement projection onto the L1, L2 and L∞ balls; verify correctness numerically
- [ ] **C**: Module 9, start: AI Evasion, First-Order Attacks
- [ ] **L**: Notebook entry, commit

### Sun 2026-10-25

- [ ] **L**: Week 5 review: projection theory, read in the same week it gets used

## Week 6: GCG from scratch

### Mon 2026-10-26

- [ ] **M**: Combinatorial optimization: the relaxation gap, LP relaxation of an integer program
- [ ] **B**: GCG: implement the loss
- [ ] **R**: GCG paper, third pass: by now every term should be familiar
- [ ] **L**: Notebook entry, commit

### Tue 2026-10-27

- [ ] **M**: Wright and Recht: coordinate descent
- [ ] **B**: GCG: top-k gradient candidate selection
- [ ] **C**: Module 9: AI Evasion, First-Order Attacks
- [ ] **L**: Notebook entry, commit

### Wed 2026-10-28

- [ ] **B**: GCG: full loop running end to end, one prompt, one model
- [ ] **R**: Andriushchenko, adaptive attacks
- [ ] **L**: Notebook entry, commit

### Thu 2026-10-29

- [ ] **B**: GCG: multi-seed runs; log every optimization curve
- [ ] **C**: Module 9, finish: AI Evasion, First-Order Attacks
- [ ] **L**: Notebook entry, commit

### Fri 2026-10-30

- [ ] **B**: GCG: ASR across a small harmful-behaviour set, with seeds and budgets recorded
- [ ] **X**: Fellowship deadline: submit if the answer was yes
- [ ] **L**: Notebook entry, commit

### Sat 2026-10-31

- [ ] **B**: Failure taxonomy: categorize every failed GCG run by apparent cause
- [ ] **C**: Module 10, start: AI Evasion, Sparsity Attacks
- [ ] **L**: Notebook entry, commit

### Sun 2026-11-01

- [ ] **L**: Week 6 review

## Week 7: Embedding-space attacks and the gap number

### Mon 2026-11-02

- [ ] **B**: Embedding-space PGD: continuous attack on input embeddings
- [ ] **R**: Schwinn on soft prompts and continuous attacks
- [ ] **L**: Notebook entry, commit

### Tue 2026-11-03

- [ ] **B**: Continuous attack: measure ASR with embeddings left unconstrained
- [ ] **C**: Module 10: AI Evasion, Sparsity Attacks
- [ ] **L**: Notebook entry, commit

### Wed 2026-11-04

- [ ] **B**: Projection back to token space: nearest neighbour in embedding space; measure what it costs
- [ ] **R**: Carlini, Are aligned neural networks adversarially aligned?
- [ ] **L**: Notebook entry, commit

### Thu 2026-11-05

- [ ] **B**: The measurement: discrete ASR against continuous ASR, same prompts, same budget
- [ ] **C**: Module 10, finish: AI Evasion, Sparsity Attacks
- [ ] **L**: Notebook entry, commit

### Fri 2026-11-06

- [ ] **B**: Repeat the gap measurement on two more models
- [ ] **L**: Notebook entry, commit

### Sat 2026-11-07

- [ ] **B**: Write up the preliminary gap numbers with bootstrap confidence intervals
- [ ] **L**: Notebook entry, commit

### Sun 2026-11-08

- [ ] **X**: Milestone: a preliminary discrete-versus-continuous gap number exists. Push it to the repo.
- [ ] **L**: Week 7 review

## Week 8: Close the COAE path, research goes light

### Mon 2026-11-09

- [ ] **C**: Module 12: AI Defense
- [ ] **B**: Refusal direction: extract it
- [ ] **L**: Notebook entry, commit

### Tue 2026-11-10

- [ ] **C**: Module 12: AI Defense
- [ ] **B**: Refusal direction: ablate it and measure the effect on ASR
- [ ] **L**: Notebook entry, commit

### Wed 2026-11-11

- [ ] **C**: Remaining modules: 2 (Applications of AI in InfoSec), 6 (AI Data Attacks) and 11 (AI Privacy), plus anything still open
- [ ] **L**: Notebook entry, commit

### Thu 2026-11-12

- [ ] **C**: Path complete. Buy the exam voucher.
- [ ] **R**: Reading only today
- [ ] **L**: Notebook entry, commit

### Fri 2026-11-13

- [ ] **C**: Re-run the two hardest labs cold, no notes
- [ ] **L**: Notebook entry, commit

### Sat 2026-11-14

- [ ] **C**: Exam prep: report template, tooling checklist, note-taking setup
- [ ] **L**: Notebook entry, commit

### Sun 2026-11-15

- [ ] **X**: Schedule the exam window. Then rest.

## Week 9: Exam week

### Mon 2026-11-16

- [ ] **C**: Exam engagement, day 1: enumeration and scoping
- [ ] **R**: Research is reading only this week

### Tue 2026-11-17

- [ ] **C**: Exam engagement, day 2
- [ ] **R**: Reading only

### Wed 2026-11-18

- [ ] **C**: Exam engagement, day 3
- [ ] **R**: Reading only

### Thu 2026-11-19

- [ ] **C**: Exam engagement, day 4: start the report as you go, not at the end
- [ ] **R**: Reading only

### Fri 2026-11-20

- [ ] **C**: Exam engagement, day 5: report body
- [ ] **R**: Reading only

### Sat 2026-11-21

- [ ] **C**: Report: findings, evidence, remediation. Treat it as a rehearsal for the research write-up.

### Sun 2026-11-22

- [ ] **X**: Submit the report.
- [ ] **L**: Week 9 review, and re-plan the last three weeks against reality

## Week 10: Eval harness and grader disagreement

### Mon 2026-11-23

- [ ] **B**: Refactor everything into one configurable eval runner
- [ ] **L**: Notebook entry, commit

### Tue 2026-11-24

- [ ] **B**: Harness: config files, pinned seeds, deterministic reruns that actually match
- [ ] **L**: Notebook entry, commit

### Wed 2026-11-25

- [ ] **B**: Implement three graders: substring match, refusal classifier, LLM judge
- [ ] **L**: Notebook entry, commit

### Thu 2026-11-26

- [ ] **B**: Grader disagreement: pairwise agreement and Cohen's kappa across the three
- [ ] **L**: Notebook entry, commit

### Fri 2026-11-27

- [ ] **B**: Hand-label 100 samples as ground truth; measure each grader's error rate against it
- [ ] **L**: Notebook entry, commit

### Sat 2026-11-28

- [ ] **B**: Re-run the gap experiment under all three graders. Does the gap survive the grader choice?
- [ ] **L**: Notebook entry, commit

### Sun 2026-11-29

- [ ] **L**: Week 10 review

## Week 11: Error bars, ablations, the base/instruct ladder

### Mon 2026-11-30

- [ ] **B**: Bootstrap confidence intervals on every headline number in the repo
- [ ] **L**: Notebook entry, commit

### Tue 2026-12-01

- [ ] **B**: Harness ablation: token budget, step budget, context truncation
- [ ] **L**: Notebook entry, commit

### Wed 2026-12-02

- [ ] **B**: Harness ablation continued: how much apparent robustness is a harness artifact?
- [ ] **L**: Notebook entry, commit

### Thu 2026-12-03

- [ ] **B**: Base versus instruct ladder: same family, run the gap experiment on both
- [ ] **L**: Notebook entry, commit

### Fri 2026-12-04

- [ ] **B**: Add a third model to the ladder
- [ ] **L**: Notebook entry, commit

### Sat 2026-12-05

- [ ] **B**: Failure attribution: classify every failure as model, search, harness or grader
- [ ] **L**: Notebook entry, commit

### Sun 2026-12-06

- [ ] **L**: Week 11 review

## Week 12: Ship it

### Mon 2026-12-07

- [ ] **B**: Repo: README, reproduce.sh, configs, every seed pinned
- [ ] **L**: Notebook entry, commit

### Tue 2026-12-08

- [ ] **B**: Write-up draft: motivation, method, results
- [ ] **L**: Notebook entry, commit

### Wed 2026-12-09

- [ ] **B**: Write-up: limitations, failure cases, proposed follow-up experiments
- [ ] **L**: Notebook entry, commit

### Thu 2026-12-10

- [ ] **B**: Figures: the gap plot, convergence curves, grader disagreement matrix
- [ ] **L**: Notebook entry, commit

### Fri 2026-12-11

- [ ] **B**: Fresh-clone test on a clean machine. Fix whatever breaks.
- [ ] **L**: Notebook entry, commit

### Sat 2026-12-12

- [ ] **X**: Publish the write-up. Start the next fellowship application, in my own words.
- [ ] **L**: Notebook entry, commit

### Sun 2026-12-13

- [ ] **X**: Send the write-up to the researchers whose work it builds on.
- [ ] **L**: Final review. Twelve weeks done.

## After the twelve weeks

### December 2026: Publish and pitch

- Publish the research write-up and the repo behind it.
- Submit talks to security conferences, starting with local BSides events.
- Scope one tool precisely before writing any code.

### January – February 2027: Build something real

- One open-source tool that tests agent systems: tool-calling abuse, permission escalation, and indirect injection into agent loops.
- Ugly and working beats elegant and unfinished. Then a README and a proper threat-model write-up.
- Release it where practitioners actually are.

### March 2027: Hunt

- Start on AI and LLM bug bounty programs, using the tool as an edge.
- Keep a weekly CTF habit and keep the tool maintained.

## Checkpoints

- **Mid-December 2026: Foundations proven.** Certification passed. Research write-up and reproducible repo published.
- **Mid-March 2027: Something people use.** Agent-security tool released with a threat-model write-up. At least one conference talk submitted.
- **Mid-June 2027: Real-world signal.** A valid finding in an AI bug bounty program, or a tool with real users.
