Aengus Lynch, Ph.D
I'm an AI safety researcher at Theorem in San Francisco, building formally verified software because I believe this is the best approach we have to steer AI during the transition to ASI.
My research has moved up the supervision stack from mechanistic interpretability, to steering vectors, to jailbreaking, to alignment evals, and now to formally verified outputs. I'm here because this is the level of supervision which requires the fewest assumptions about what we can monitor in AI systems.
I'm the first author of the Anthropic blackmail study, Agentic Misalignment, research which showed frontier models engaging in blackmail when pursuing goals, covered by dozens of major outlets and quoted in the US Congress. That research led me to found Watertight AI, a third-party auditing startup selling alignment evals for coding-agent deployment; I shut it down a year later when it became clear the market didn't need it yet. I subsequently first-authored a 2026 follow-up, Agentic Misalignment in Summer 2026. I hold a PhD in Artificial Intelligence from University College London, having published the thesis "The Persistent Vulnerability of Aligned AI Systems."
Watch
Gemini CLI — a production coding agent — drafting coercive emails in the blackmail scenario, first try.
The Bureau of Investigative Journalism: I walk through the blackmail research — and re-run it live on today's models.
Research
Agentic Misalignment in Summer 2026 (2026)
Aengus Lynch, John Hughes, Alex Serrano, Robert Kirk, Samuel R. Bowman
Anthropic Alignment Science Blog, July 2026. Documents four failure modes in frontier models: covert sabotage, assisting fraud, motivated mislabeling, and coaching human proxies to whistleblow.
PhD Thesis: The Persistent Vulnerability of Aligned AI Systems (2025)
Aengus Lynch
Supervised by Ricardo Silva. Examined by Florian Tramèr and Ilija Bogunovic.
Agentic Misalignment: How LLMs Could be Insider Threats (2025)
Aengus Lynch, Benjamin Wright, Caleb Larson, Kevin K. Troy, Stuart J. Ritchie, Sören Mindermann, Ethan Perez, Evan Hubinger
Demonstrated that frontier models from major AI labs will engage in blackmail, deception, and harmful behaviors when pursuing goals. Featured in the Claude 4 system card.
Best-of-N Jailbreaking (2024)
John Hughes*, Sara Price*, Aengus Lynch*, Rylan Schaeffer, Fazl Barez, Sanmi Koyejo, Henry Sleight, Erik Jones, Ethan Perez, Mrinank Sharma
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs (2024)
Abhay Sheshadri*, Aidan Ewart*, Phillip Guo*, Aengus Lynch*, Cindy Wu*, Vivek Hebbar*, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, Stephen Casper
Analyzing the generalization and reliability of steering vectors (2024)
Daniel Tan, David Chanin, Aengus Lynch, Adrià Garriga-Alonso, Brooks Paige, Dimitrios Kanoulas, Robert Kirk
Eight methods to evaluate robust unlearning in LLMs (2024)
Aengus Lynch*, Phillip Guo*, Aidan Ewart*, Stephen Casper, Dylan Hadfield-Menell
Towards automated circuit discovery for mechanistic interpretability (2023)
Arthur Conmy*, Augustine N. Mavor-Parker*, Aengus Lynch*, Stefan Heimersheim, Adrià Garriga-Alonso
Spotlight at NeurIPS 2023
Spawrious: A benchmark for fine control of spurious correlation biases (2023)
Aengus Lynch*, Gbètondji J-S Dovonon*, Jean Kaddour*, Ricardo Silva
Causal machine learning: A survey and open problems (2022)
Jean Kaddour*, Aengus Lynch*, Qi Liu, Matt J. Kusner, Ricardo Silva