Projects
Things I built
Research I'm running, tools I've built, and security work from before the AI pivot. Every project here has a clear question, a method, and numbers I can defend, or it says plainly that it isn't finished yet.
-
Attack Headroom#
When gradient-based adversarial attacks fail against safety-tuned language models, how much of that failure is real robustness, and how much is the search, the harness, or the grader?
A failed jailbreak isn't evidence that no vulnerability exists. This project runs the same harmful behaviours under progressively stronger attacker access, and attributes every failure to a specific cause.
- Attacker access ladder: black-box → discrete gradient-guided (GCG) → continuous embedding-space → activation intervention.
- Every failure gets one label: model robustness, optimization failure, harness failure, grader failure, refusal-detector failure, context/budget artifact, or continuous-to-discrete projection loss.
- Three independent graders (string match, a guard classifier, the HarmBench classifier). Their disagreement is reported as a result in its own right.
- Benign control sets (JailbreakBench benign behaviours and XSTest) separate "the attack failed" from "the grader mislabelled a harmless answer".
- Difficulty ladder from a model with no safety training up to Llama-2-chat, plus base/instruct ablation pairs where safety tuning is the only difference.
Research rig built and verified: 12/12 acceptance tests, including the gradient with respect to a one-hot token representation (GCG's mechanical core) over a 151,936-token vocabulary at 2.5 GB peak VRAM. Now in Phase 0: building GPT from scratch before attacking anything.
- PyTorch
- Transformers
- TransformerLens
- nnsight
- HarmBench
- JailbreakBench
- XSTest
-
Where refusals actually come from#
How does the framing of a legitimate security request move a model's refusal boundary?
780 probes across four local open-weight models, on legitimate security work such as exploit education, malware reverse engineering, authorised offensive testing, CTFs and vulnerability research.
- Framing effects ran backwards from intuition. "Professional" and "academic" framings produced zero friction. Every hard refusal, and 7 of 9 hedges, came from "hypothetical" ("for a novel") and "urgency" ("we're mid-incident") framings.
- Models named the urgency framing explicitly, and then declined. Claimed time pressure made them more cautious, not less.
- The standard benign-vs-harmful suite scored 480/480 on every model and couldn't tell them apart. Only a hard-benign suite separated them.
- Two grader traps each produced a false headline. First, reasoning models spent their whole token budget thinking and returned empty content, which got scored as compliance. Second, a regex matched "I cannot" inside "While I cannot…", so hedged help got scored as refusal.
- The fix was a three-way REFUSED / HEDGED / COMPLIED label, with full responses stored so every run can be re-graded offline.
- llama.cpp
- Python
- Gemma
- Qwen
- Deterministic grading
-
Local LLM evaluation rig#
Which open-weight models are actually useful for security work on a single 16 GB GPU, and how do you measure that honestly?
A deterministic benchmark harness for open-weight models: a 42-item cybersecurity suite across 14 categories, plus a 27-item general suite, graded with no LLM judge.
- Security categories include CVE/CVSS reasoning, secure code review, exploit reasoning, log triage, detection engineering, ATT&CK mapping, network forensics, malware triage, cloud/IaC, web appsec, binary exploitation and refusal calibration.
- Grading is fully deterministic: executed Python, parsed JSON, weighted checks and needle-in-a-haystack recall.
- Every model passes a coherence gate before its scored run. A corrupted quantisation looks like "quality degradation", and scoring it would be meaningless.
- Greedy decoding for every model so comparisons are apples-to-apples. Reasoning models get a larger token budget, or their chain of thought eats the answer.
- llama.cpp
- CUDA
- Python
- RTX 5070 Ti
-
AI-assisted CVE analysis#
Can LLMs triage the incoming stream of CVEs against a real system's architecture?
An LLM and graph-database pipeline that maps incoming NIST CVE data onto a system's architecture and call trees, and automates triage and enrichment.
- Extracts exploit techniques, affected components and likely attack paths from unstructured vulnerability data.
- Maps findings onto architecture and call-tree graphs in Neo4j to prioritise what actually matters.
- Built as part of security engineering work in aviation. No code or data is public.
- OpenAI models
- Neo4j
- MITRE ATT&CK
-
Android malware analysis#
What do malicious Android apps actually do once they're running, and how do you detect it?
Led a student research team at UW–Whitewater, under Dr. Chandra Sharma, doing static and dynamic analysis of malicious APKs to develop detection methods.
- Sandboxed samples in Genymotion and observed their live behaviour and network traffic.
- Used INetSim to simulate command-and-control servers, intercepting traffic for deeper network analysis.
- Genymotion
- INetSim
- APKTool
- Ghidra
-
This website#
What's the least friction between finishing a day of study and publishing it?
A static site generator and a private local writing studio, built from scratch and deployed to Cloudflare. No CMS, no trackers, no database.
- Markdown with math (KaTeX), code highlighting, callouts and footnotes, rendered at build time.
- A local-only studio with live preview, image optimisation and one-click publish.
- Auto-generated social cards, RSS, a sitemap, and a roadmap that updates itself every day.
- Node.js
- Cloudflare Pages
- KaTeX
- Playwright