I'm a security engineer, and I'm moving into AI safety and security. Rather than doing that quietly, I'm doing it here, in public, one logged day at a time.
Why "safety in the layers"#
In security we talk about defense in depth. No single control is perfect, so you stack them, and the system is safe because of the layers, not despite them.
Language models are layers too. Whatever makes a model refuse, comply, or get talked out of refusing lives somewhere in the stack, in the residual stream and in the directions it writes to. Getting good at AI safety means understanding both kinds of layers: the controls we wrap around models, and the machinery inside them.
The logo shows a model picking its next token. After "safety in the", the candidates are layers at 0.91, with jailbreak and breach struck out. That's the goal, in one picture.
The plan#
For the next 180 days I'm locked in on a public roadmap:
- Math, just in time. Linear algebra, calculus, probability, statistics and convex optimization, each learned the week before it's needed.
- Everything from scratch. micrograd, nanoGPT, attention and KV caching, then FGSM, PGD and GCG. I want to understand what I'm about to attack.
- ARENA for transformer internals and interpretability.
- Hack The Box's AI red-teaming certification, as the structured offensive track.
- A real research question: when an adversarial attack fails against a safety-tuned model, how much of that is genuine robustness, and how much is search, harness or grader failure?
After the first twelve weeks, the plan is to publish that research, build an open-source tool for testing agent systems, and start hunting in AI bug bounty programs.
What I'll post#
Everything is written by me, and every number comes with its method. When I'm wrong, the correction goes in the log too.
If you're making the same move, from security or anywhere else, the Start here page is the path I'd hand a friend. Follow along on Instagram for the daily version.
Let's go.