About
Hi, I’m Prathik
Hi, I'm Prathik.
I'm a security engineer from Wisconsin. My background is offensive security: malware analysis, reverse engineering, red teaming, and most recently security testing in aviation. I've spent my time learning how systems break. Now I'm pointing that at the systems I think matter most: AI models, and the agents we're starting to wire into everything.
Safety in the Layers is my public notebook for that shift. The name works two ways. In security, safety lives in layers. That's defense in depth: no single control has to be perfect. In a neural network, behaviour lives in layers too. Refusals, representations and failure modes all sit somewhere in the residual stream. I want to understand both, and I want to be able to attack a model and explain why the attack worked.
What I'm doing right now#
I'm locked in on a structured ramp-up. The full roadmap is public. It covers math from linear algebra to convex optimization, building GPT and GCG from scratch, the ARENA curriculum, Hack The Box's AI red-teaming certification, and a research project on why adversarial attacks fail. I'm also applying to BlueDot Impact's AI safety courses.
The question I keep coming back to: when an attack fails against a safety-tuned model, how much of that is real robustness, and how much is a weak search, a leaky harness, or a bad grader?
What I post#
- The Log: what I studied, built and broke, day by day.
- Paper Notes: research papers broken down: the claim, the method, the numbers, and what I think.
- Builds & Research: the harnesses, attacks and evals I'm building, with results.
- Explainers: concepts explained the way I wish they'd been explained to me.
I'm not an expert yet. This is where I work in the open until I am. If you're on a similar path, or you're further along and see me doing something dumb, I'd love to hear from you.
Now updated September 28, 2026
Week 2 of the lock-in. The math block that everything else depends on: projections and least squares, Gram–Schmidt, eigenvalues, the SVD, norms, and just enough multivariable calculus to derive backprop by hand.
Building: Karpathy's micrograd from scratch, then makemore parts 1–5, then Let's Build GPT start to finish. No autograd I didn't write myself this week.
Reading: the GCG paper (Universal and Transferable Adversarial Attacks on Aligned Language Models), with a full first pass and a second pass focused only on the loss. Also A Mathematical Framework for Transformer Circuits, parts 1 and 2, plus the rest of Jailbroken.
Certification: Hack The Box's AI red-teaming path, one module at a time.
Applying: BlueDot Impact's AI safety course.
End-of-week test: projections, SVD and gradients, from memory, on paper.
This page is a now page. It says what I'm focused on at the moment, and I update it as that changes.
Timeline
-
2026 –
Learning AI safety & security in public, Safety in the Layers
A structured ramp-up into alignment, interpretability and adversarial robustness, logged daily.
-
2025 –
Data Security Engineer Intern
Security engineering in the aviation industry, including an AI-assisted CVE analysis pipeline.
-
2025
Founder, LLM-based adversarial testing, Microsoft for Startups Founders Hub
Automated, containerised environments for testing and analysing adversarial behaviour.
-
2024–25
Lead Researcher, Android malware analysis, UW–Whitewater
Static and dynamic analysis of malicious APKs, with Dr. Chandra Sharma.
-
2024–25
Cyber Security Intern, Malware Analyst, Cyber Crime Department, Coimbatore
Incident response for 20+ daily complaints: device compromise, fraud, and malware removal.
Certifications
- HTB Certified Offensive AI Expert (COAE)
- Azure Security Engineer Associate
- Security+
Education
- University of Wisconsin–Whitewater
Competitions
- Palo Alto Networks Secure the Future: finalist (finance-sector research)
- picoCTF · NCL · CCDC · CyberPatriot
Say hello
The fastest way to reach me is a DM on Instagram or LinkedIn. I'm especially happy to hear from people making the same move into AI safety, and from anyone further along who's willing to point out what I'm missing.