Roadmap
The roadmap
This is the plan I'm actually following. It isn't a sanitised version for the internet. It's the same list I tick off every day. Twelve detailed weeks of math, builds, certification and research, then the next three months in outline. If you're starting out, steal it.
Six tracks
How the days are built
- M
Math
Linear algebra, calculus, probability, statistics and optimization, learned the week before they're needed.
- B
Build
Everything from scratch: micrograd, nanoGPT, FGSM, PGD, GCG, eval harnesses.
- C
COAE
Hack The Box’s AI Red Teamer path, built with Google, and its practical exam, the HTB Certified Offensive AI Expert: seven days in a live lab, ending in a professional report.
- R
Read
Papers and tutorials, often in several passes.
- L
Log
A lab-notebook entry and a commit, every single day.
- X
Milestone
Checkpoints that prove the work happened.
The first twelve weeks
Five phases
-
01 · Weeks 1–4
Foundations
Linear algebra, calculus, probability and statistics, alongside Karpathy's builds, nanoGPT from scratch, and ARENA chapters 0 and 1.
-
02 · Weeks 5–7
Attacks from scratch
Convex optimization, then FGSM, PGD, GCG and embedding-space attacks, ending in a measured discrete-versus-continuous gap.
-
03 · Weeks 8–9
Certification
Close out the COAE path and sit the multi-day practical exam. Research drops to reading only.
-
04 · Weeks 10–11
Measurement
One eval harness, three graders, hand-labelled ground truth, bootstrap error bars and harness ablations.
-
05 · Weeks 12
Ship
Clean repo, reproducible runs, figures, and a public write-up.
Day by day
Every week, every task
W1 Linear algebra sprint, COAE speedrun, environment Done
-
Tue Sep 22
- M3Blue1Brown, Essence of Linear Algebra: all 16 videos in one sitting
- BEnvironment: CUDA 12.8, torch cu128, verify the 5070 Ti is visible
- BWrite check_env.py; repo skeleton with src/, experiments/, configs/, .gitignore
- LStart the lab notebook file. First commit.
-
Wed Sep 23
- MStrang 18.06 lectures 1–3 (vectors, elimination, matrix multiplication) + problem set
- MNumPy from scratch: matrix multiply, Gaussian elimination
- CModule 1: Fundamentals of AI
- BChat-template forensics: dump raw templates for three models, diff them
- RGCG paper: abstract and intro only, orientation pass
- LNotebook entry, commit
-
Thu Sep 24
- MStrang lectures 4–6 (LU, transpose, vector spaces) + problem set
- CModule 3: Introduction to Red Teaming AI
- BTokenizer exploration: BPE merges, glitch tokens, whitespace behaviour across three tokenizers
- RKarpathy, Let's Build the GPT Tokenizer
- LNotebook entry, commit
-
Fri Sep 25
- MStrang lectures 7–9 (nullspace, Ax=b, independence and basis) + problem set
- MNumPy: compute rank and nullspace by hand, check against np.linalg
- CModule 4: Prompt Injection Attacks
- BLogprob harness: per-token logprobs for a prompt set, written to JSONL
- RJailbroken (Wei et al.), sections 1–3
- LNotebook entry, commit
-
Sat Sep 26
- MStrang lectures 10–12 (four subspaces, orthogonality) + problem set
- CModule 5: LLM Output Attacks
- BExtend the harness: refusal-token probability tracker
- LNotebook entry, commit
-
Sun Sep 27
- MOn paper, from memory: derive projection onto a subspace
- LWeek 1 review: what stuck, what didn't, what blocked me
- XFellowship applications: confirm deadlines and put a decision date on the calendar
W2 Projections, eigen-structure, SVD, calculus, Karpathy This week
-
Mon Sep 28
- MStrang lectures 14–16 (projections, least squares); implement the projection matrix in NumPy
- CModule 7: Attacking AI, Application and System
- BRepo hygiene: config files, seed handling, reproduce.sh stub
- RA Mathematical Framework for Transformer Circuits, part 1
- LNotebook entry, commit
-
Tue Sep 29
- MStrang lectures 17–19 (Gram-Schmidt, QR, determinants)
- BKarpathy micrograd, built from scratch, start to finish
- RTransformer Circuits part 2: attention heads, QK and OV circuits
- LNotebook entry, commit
-
Wed Sep 30
- MStrang lectures 21–22 (eigenvalues, diagonalization); power iteration in NumPy
- Bmakemore parts 1 and 2
- RGCG paper, full first pass
- LNotebook entry, commit
-
Thu Oct 1
- MStrang lectures 29–30 (SVD); implement low-rank approximation from scratch
- Bmakemore parts 3 and 4
- RGCG paper, second pass: take notes specifically on the loss
- LNotebook entry, commit
-
Fri Oct 2
- MNorms: L0, L1, L2, L∞. Implement projection onto each norm ball in NumPy
- Bmakemore part 5
- RJailbroken, remainder
- LNotebook entry, commit
-
Sat Oct 3
- MKhan multivariable: partials, gradient, directional derivative, chain rule, Jacobian, Taylor. Stop there.
- BLet's Build GPT, start to finish
- RKolter and Madry tutorial, chapters 1–3
- LNotebook entry, commit
-
Sun Oct 4
- MOn paper: derive backprop for a two-layer MLP
- XSelf-test: projections, SVD, gradients. The blocking math block ends here.
- LWeek 2 review
Written this week
W3 Probability and nanoGPT Upcoming
-
Mon Oct 5
- MBlitzstein Stat 110, lectures 1–4: counting, conditional probability, Bayes
- BnanoGPT from scratch, day 1
- RAttention Is All You Need, close read
- LNotebook entry, commit
-
Tue Oct 6
- MBlitzstein 5–8: random variables, expectation, the named distributions
- BnanoGPT day 2: train on a tiny corpus, loss curve looks sane
- BARENA Chapter 0 setup
- LNotebook entry, commit
-
Wed Oct 7
- MBlitzstein 9–11: conditional expectation
- BAttention implemented by hand: no nn.MultiheadAttention anywhere
- BARENA Chapter 0 exercises
- LNotebook entry, commit
-
Thu Oct 8
- MBlitzstein 12–14; derive MLE for softmax and cross-entropy by hand
- BKV cache from scratch; assert equivalence against the uncached path
- BARENA Chapter 0
- LNotebook entry, commit
-
Fri Oct 9
- MSampling: implement temperature, top-k and top-p from raw logits
- BForward hooks: extract the residual stream at every layer
- RThe logit lens
- LNotebook entry, commit
-
Sat Oct 10
- XFellowship decision: decide, and start the application if the answer is yes
- BLogit lens implemented across all layers of one model
- LNotebook entry, commit
-
Sun Oct 11
- LWeek 3 review: record the fellowship decision and the reason for it
W4 Statistical inference, LoRA, ARENA 1 Upcoming
-
Mon Oct 12
- MWasserman, All of Statistics, chapters 6–7
- BLoRA fine-tune of a small instruct model
- BARENA Chapter 1, start
- LNotebook entry, commit
-
Tue Oct 13
- MWasserman chapter 8; implement a bootstrap confidence interval in NumPy
- BLoRA continued: evaluate before and after, same prompts
- RRefusal in LLMs is mediated by a single direction
- LNotebook entry, commit
-
Wed Oct 14
- MWasserman chapters 9–10: hypothesis testing
- BARENA Chapter 1, attention exercises
- LNotebook entry, commit
-
Thu Oct 15
- MThe high-leverage block: Wilson interval, Clopper-Pearson, rule of three. Implement all three.
- MPlot coverage for each, small n. See where the normal approximation lies to you.
- BARENA Chapter 1
- LNotebook entry, commit
-
Fri Oct 16
- MWrite the note: what does 0 successes in 200 attempts actually license you to claim?
- BARENA Chapter 1, finish
- LNotebook entry, commit
-
Sat Oct 17
- BSlack day: clear whatever is behind
- BIntegrate nanoGPT, hooks and harness into one clean module
- LNotebook entry, commit
-
Sun Oct 18
- XFoundations phase complete. Everything after this is the project.
- LWeek 4 review
W5 Convex optimization, COAE 8, FGSM and PGD Upcoming
-
Mon Oct 19
- MBoyd EE364A lectures 1–3: convex sets
- CModule 8, start: AI Evasion, Foundations
- RMadry et al., Towards Deep Learning Models Resistant to Adversarial Attacks
- LNotebook entry, commit
-
Tue Oct 20
- MBoyd lectures 4–5: convex functions
- BFGSM from scratch against a small vision model
- LNotebook entry, commit
-
Wed Oct 21
- MBoyd lectures 6–7
- BPGD from scratch, both L∞ and L2
- CModule 8: AI Evasion, Foundations
- LNotebook entry, commit
-
Thu Oct 22
- MWright and Recht chapter 3: gradient descent and convergence rates
- BPGD step-size sweep; save the convergence curves
- LNotebook entry, commit
-
Fri Oct 23
- MWright and Recht chapter 4
- CModule 8, finish: AI Evasion, Foundations
- RBackground for the malicious fine-tuning direction
- LNotebook entry, commit
-
Sat Oct 24
- MImplement projection onto the L1, L2 and L∞ balls; verify correctness numerically
- CModule 9, start: AI Evasion, First-Order Attacks
- LNotebook entry, commit
-
Sun Oct 25
- LWeek 5 review: projection theory, read in the same week it gets used
W6 GCG from scratch Upcoming
-
Mon Oct 26
- MCombinatorial optimization: the relaxation gap, LP relaxation of an integer program
- BGCG: implement the loss
- RGCG paper, third pass: by now every term should be familiar
- LNotebook entry, commit
-
Tue Oct 27
- MWright and Recht: coordinate descent
- BGCG: top-k gradient candidate selection
- CModule 9: AI Evasion, First-Order Attacks
- LNotebook entry, commit
-
Wed Oct 28
- BGCG: full loop running end to end, one prompt, one model
- RAndriushchenko, adaptive attacks
- LNotebook entry, commit
-
Thu Oct 29
- BGCG: multi-seed runs; log every optimization curve
- CModule 9, finish: AI Evasion, First-Order Attacks
- LNotebook entry, commit
-
Fri Oct 30
- BGCG: ASR across a small harmful-behaviour set, with seeds and budgets recorded
- XFellowship deadline: submit if the answer was yes
- LNotebook entry, commit
-
Sat Oct 31
- BFailure taxonomy: categorize every failed GCG run by apparent cause
- CModule 10, start: AI Evasion, Sparsity Attacks
- LNotebook entry, commit
-
Sun Nov 1
- LWeek 6 review
W7 Embedding-space attacks and the gap number Upcoming
-
Mon Nov 2
- BEmbedding-space PGD: continuous attack on input embeddings
- RSchwinn on soft prompts and continuous attacks
- LNotebook entry, commit
-
Tue Nov 3
- BContinuous attack: measure ASR with embeddings left unconstrained
- CModule 10: AI Evasion, Sparsity Attacks
- LNotebook entry, commit
-
Wed Nov 4
- BProjection back to token space: nearest neighbour in embedding space; measure what it costs
- RCarlini, Are aligned neural networks adversarially aligned?
- LNotebook entry, commit
-
Thu Nov 5
- BThe measurement: discrete ASR against continuous ASR, same prompts, same budget
- CModule 10, finish: AI Evasion, Sparsity Attacks
- LNotebook entry, commit
-
Fri Nov 6
- BRepeat the gap measurement on two more models
- LNotebook entry, commit
-
Sat Nov 7
- BWrite up the preliminary gap numbers with bootstrap confidence intervals
- LNotebook entry, commit
-
Sun Nov 8
- XMilestone: a preliminary discrete-versus-continuous gap number exists. Push it to the repo.
- LWeek 7 review
W8 Close the COAE path, research goes light Upcoming
-
Mon Nov 9
- CModule 12: AI Defense
- BRefusal direction: extract it
- LNotebook entry, commit
-
Tue Nov 10
- CModule 12: AI Defense
- BRefusal direction: ablate it and measure the effect on ASR
- LNotebook entry, commit
-
Wed Nov 11
- CRemaining modules: 2 (Applications of AI in InfoSec), 6 (AI Data Attacks) and 11 (AI Privacy), plus anything still open
- LNotebook entry, commit
-
Thu Nov 12
- CPath complete. Buy the exam voucher.
- RReading only today
- LNotebook entry, commit
-
Fri Nov 13
- CRe-run the two hardest labs cold, no notes
- LNotebook entry, commit
-
Sat Nov 14
- CExam prep: report template, tooling checklist, note-taking setup
- LNotebook entry, commit
-
Sun Nov 15
- XSchedule the exam window. Then rest.
W9 Exam week Upcoming
-
Mon Nov 16
- CExam engagement, day 1: enumeration and scoping
- RResearch is reading only this week
-
Tue Nov 17
- CExam engagement, day 2
- RReading only
-
Wed Nov 18
- CExam engagement, day 3
- RReading only
-
Thu Nov 19
- CExam engagement, day 4: start the report as you go, not at the end
- RReading only
-
Fri Nov 20
- CExam engagement, day 5: report body
- RReading only
-
Sat Nov 21
- CReport: findings, evidence, remediation. Treat it as a rehearsal for the research write-up.
-
Sun Nov 22
- XSubmit the report.
- LWeek 9 review, and re-plan the last three weeks against reality
W10 Eval harness and grader disagreement Upcoming
-
Mon Nov 23
- BRefactor everything into one configurable eval runner
- LNotebook entry, commit
-
Tue Nov 24
- BHarness: config files, pinned seeds, deterministic reruns that actually match
- LNotebook entry, commit
-
Wed Nov 25
- BImplement three graders: substring match, refusal classifier, LLM judge
- LNotebook entry, commit
-
Thu Nov 26
- BGrader disagreement: pairwise agreement and Cohen's kappa across the three
- LNotebook entry, commit
-
Fri Nov 27
- BHand-label 100 samples as ground truth; measure each grader's error rate against it
- LNotebook entry, commit
-
Sat Nov 28
- BRe-run the gap experiment under all three graders. Does the gap survive the grader choice?
- LNotebook entry, commit
-
Sun Nov 29
- LWeek 10 review
W11 Error bars, ablations, the base/instruct ladder Upcoming
-
Mon Nov 30
- BBootstrap confidence intervals on every headline number in the repo
- LNotebook entry, commit
-
Tue Dec 1
- BHarness ablation: token budget, step budget, context truncation
- LNotebook entry, commit
-
Wed Dec 2
- BHarness ablation continued: how much apparent robustness is a harness artifact?
- LNotebook entry, commit
-
Thu Dec 3
- BBase versus instruct ladder: same family, run the gap experiment on both
- LNotebook entry, commit
-
Fri Dec 4
- BAdd a third model to the ladder
- LNotebook entry, commit
-
Sat Dec 5
- BFailure attribution: classify every failure as model, search, harness or grader
- LNotebook entry, commit
-
Sun Dec 6
- LWeek 11 review
W12 Ship it Upcoming
-
Mon Dec 7
- BRepo: README, reproduce.sh, configs, every seed pinned
- LNotebook entry, commit
-
Tue Dec 8
- BWrite-up draft: motivation, method, results
- LNotebook entry, commit
-
Wed Dec 9
- BWrite-up: limitations, failure cases, proposed follow-up experiments
- LNotebook entry, commit
-
Thu Dec 10
- BFigures: the gap plot, convergence curves, grader disagreement matrix
- LNotebook entry, commit
-
Fri Dec 11
- BFresh-clone test on a clean machine. Fix whatever breaks.
- LNotebook entry, commit
-
Sat Dec 12
- XPublish the write-up. Start the next fellowship application, in my own words.
- LNotebook entry, commit
-
Sun Dec 13
- XSend the write-up to the researchers whose work it builds on.
- LFinal review. Twelve weeks done.
Months four to six
After the twelve weeks
Less detail, on purpose. The first twelve weeks will change what makes sense here.
Checkpoints
How I'll know it's working
-
Mid-December 2026
Foundations proven
Certification passed. Research write-up and reproducible repo published.
-
Mid-March 2027
Something people use
Agent-security tool released with a threat-model write-up. At least one conference talk submitted.
-
Mid-June 2027
Real-world signal
A valid finding in an AI bug bounty program, or a tool with real users.
Rules I'm holding myself to
Principles
- Build it from scratch before using the library version.
- Learn the math the week before it's needed, not in the abstract.
- Log every day, even when it's two lines.
- Every headline number gets an error bar.
- A failed attack isn't evidence of robustness until the harness and grader have been ruled out.
Want to take this path yourself?
The roadmap is mine. The guide for choosing your own is on the Start here page.