Aengus Lynch, Ph.D
AI Alignment Research
I am a generalist operator and AI safety researcher, currently building formally verified software environments as RL training data for frontier AI models at Theorem in San Francisco. I am first author of Anthropic's blackmail study — agentic misalignment research showing frontier models can engage in blackmail and deception when pursuing goals, covered by over 15 major outlets including BBC and Fortune — and of its 2026 follow-up, Agentic Misalignment in Summer 2026, written during the Anthropic Fellows program. I hold a PhD in AI from UCL on the persistent vulnerability of aligned AI systems.
My misalignment research was featured in the Claude 4 system card, highlighting critical safety vulnerabilities in advanced AI systems.
Recent Coverage
See more coverage →
Research
Agentic Misalignment in Summer 2026 (2026)
Aengus Lynch, John Hughes, Alex Serrano, Robert Kirk, Samuel R. Bowman
Anthropic Alignment Science Blog, July 2026. Documents four failure modes in frontier models: covert sabotage, assisting fraud, motivated mislabeling, and coaching human proxies to whistleblow.
PhD Thesis: The Persistent Vulnerability of Aligned AI Systems (2025)
Aengus Lynch
Supervised by Ricardo Silva. Examined by Florian Tramèr and Ilija Bogunovic.
Agentic Misalignment: How LLMs Could be Insider Threats (2025)
Aengus Lynch, Benjamin Wright, Caleb Larson, Kevin K. Troy, Stuart J. Ritchie, Sören Mindermann, Ethan Perez, Evan Hubinger
Demonstrated that frontier models from major AI labs will engage in blackmail, deception, and harmful behaviors when pursuing goals. Featured in the Claude 4 system card.
Best-of-N Jailbreaking (2024)
John Hughes*, Sara Price*, Aengus Lynch*, Rylan Schaeffer, Fazl Barez, Sanmi Koyejo, Henry Sleight, Erik Jones, Ethan Perez, Mrinank Sharma
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs (2024)
Abhay Sheshadri*, Aidan Ewart*, Phillip Guo*, Aengus Lynch*, Cindy Wu*, Vivek Hebbar*, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, Stephen Casper
Analyzing the generalization and reliability of steering vectors (2024)
Daniel Tan, David Chanin, Aengus Lynch, Adrià Garriga-Alonso, Brooks Paige, Dimitrios Kanoulas, Robert Kirk
Eight methods to evaluate robust unlearning in LLMs (2024)
Aengus Lynch*, Phillip Guo*, Aidan Ewart*, Stephen Casper, Dylan Hadfield-Menell
Towards automated circuit discovery for mechanistic interpretability (2023)
Arthur Conmy*, Augustine N. Mavor-Parker*, Aengus Lynch*, Stefan Heimersheim, Adrià Garriga-Alonso
Spotlight at NeurIPS 2023
Spawrious: A benchmark for fine control of spurious correlation biases (2023)
Aengus Lynch*, Gbètondji J-S Dovonon*, Jean Kaddour*, Ricardo Silva
Causal machine learning: A survey and open problems (2022)
Jean Kaddour*, Aengus Lynch*, Qi Liu, Matt J. Kusner, Ricardo Silva