professional skeptic of impressive-looking benchmark results
reward-hacking-gym — A Q-learning agent finds a reward exploit I didn't design: 27.5 mean reward, 0% true task success. A held-out audit independent of the training code catches it and gates CI. Repaired via potential-based shaping (Ng et al., 1999) — audit passes at 1.00.
sovereign-rag-ratchet — A self-improving RAG pipeline for air-gapped deployment whose optimizer never sees the evaluator. Caught a planted prompt-injection cheat (+0.33 visible accuracy, −0.33 held-out) and auto-reverted it.
cot-scratchpad-research — Direct vs. Chain-of-Thought vs. Scratchpad on a 0.79M-parameter transformer. The scratchpad model invented its own carry notation unprompted — then failed out of distribution exactly where the others did.
watt-bench — Discrete-event simulation benchmark treating rack-level power budgets as hard constraints and tokens-per-watt as a first-class metric. On real Azure traces: 1,000+ thermal throttle events for latency-greedy scheduling, zero for power-aware.

