Harry Mayne

PhD @ University of Oxford
X logoLinkedIn logo
Image of me!
I'm a PhD researcher in Oxford's OxRML and BOLD labs.
My research focuses on LLM self-explanations: the natural language explanations LLMs give to justify their own decision-making. I measure whether these explanations are faithful to a model's true internal reasoning, and I develop new training objectives to incentivise faithfulness. I'm motivated by AI safety and think monitoring verbalised reasoning is one of the most promising short-to-medium-term safety methods. I also work on safety and capability evals, including the LingOly and LingOly-TOO reasoning benchmarks, and the Measuring what Matters review paper. My work has been published at NeurIPS, ICML, and ICLR. I'm supervised by Prof. Adam Mahdi (OxRML) and Prof. Jakob Foerster (BOLD).

Previously, I was an Astra Fellow with Owain Evans (Truthful AI) and completed the SPAR programme with Noah Siegel (Google DeepMind).
Selected Publications

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

J. Betley, J. Treutlein, J. Dubiński, H. Mayne, K. Gałązka, N. Warncke, A. Sztyber-Betley, O. Evans

A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

H. Mayne*, J. Kang*, D. Gould, K. Ramchandran, A. Mahdi, N. Siegel
ICML 2026

LINGOLY-TOO: Disentangling Memorisation from Reasoning with Linguistic Templatisation and Orthographic Obfuscation

J. Khouja, K. Korgul, S. Hellsten, L. Yang, V. Neacsu, H. Mayne, R. O. Kearns, A. M. Bean, A. Mahdi
ICLR 2026

LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations

H. Mayne, R. O. Kearns, Y. Yang, A. M. Bean, E. Delaney, C. Russell, A. Mahdi
EMNLP 2025

Toxic Neurons Aren't Enough to Explain DPO: A Mechanistic Analysis for Toxicity Reduction

Y. Yang, F. Sondej, H. Mayne, A. Mahdi
EMNLP 2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

A. M. Bean, R. O. Kearns, A. Romanou, F. S. Hafner, H. Mayne, et al.
NeurIPS 2025

LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages

A. M. Bean, S. Hellsten, H. Mayne, J. Magomere, E. A. Chi, R. Chi, S. A. Hale, H. R. Kirk
NeurIPS 2024Oral · Top 0.5%

Can Sparse Autoencoders Be Used to Decompose and Interpret Steering Vectors?

H. Mayne, Y. Yang, A. Mahdi
NeurIPS 2024Workshops

Positions

Jan 2026 – Jul 2026
Astra Fellow with Owain Evans (Truthful AI)LLM generalisation during finetuning. Based at Constellation, Berkeley.
Sep 2025 – Apr 2026
SPAR with Noah Siegel (Google DeepMind)Developed new explanatory faithfulness metrics.
Jun 2025 – Jul 2026
AI Advisor, International Growth CentreUsing AI to aid public service delivery in developing countries.

Writing

I mainly write about AI safety, explainability, and evals.
Feb 2026
A Positive Case for FaithfulnessSummarising our paper on whether LLM self-explanations help predict model behaviour.
Sep 2025
LLMs Don't Know Their Own Decision BoundariesSummarising our EMNLP 2025 paper on self-generated counterfactual explanations.
Sep 2025
New to AI?A curated reading list for getting started with AI and machine learning.
Aug 2025
AI Safety Researchers Should Care About Eval QualityWhy evaluation methodology matters for AI safety research.
Mar 2025
Are Recent LLMs Better at Reasoning, or Better at Memorising?Summarising the LingOly-TOO (ICLR 2026) benchmark for disentangling memorisation from reasoning.
Dec 2024
University of Cambridge Economics Interview QuestionsA collection of practice interview questions for prospective Cambridge economics applicants.
I'm a final-year PhD researcher. I've had an unusual path, having originally studied economics.
Education
2023 –
University of OxfordDPhil Social Data Science · LLM explainability and interpretability
2022 – 2023
University of OxfordMSc Social Data Science · Distinction (77%). OII Thesis Prize
2019 – 2022
University of CambridgeBA Economics · Double First Class Honours. Patrick Cross Prize
Grants & Awards
2022 – 2027
Grand Union DTP, Economic and Social Research CouncilFull PhD Scholarship (MSc + DPhil)
2025 – 2026
Dieter Schwarz FoundationResearch agenda sponsorship
Teaching
I've held several teaching positions including TA-ing the Social Data Science MSc at Oxford and tutoring Stanford computer science students. Previous students have gone on to the CS Master's at Stanford and various PhD positions at Oxford.
2023 – 2025
Stanford UniversityMachine Learning · Personalised ML and AI tutorials for Stanford CS undergraduates on their semester abroad.
2023 – 2024
University of OxfordApplied Analytical Statistics · Teaching Assistant for the Social Data Science MSc.
2024
Oxmedica / MawhibaAI and Big Data · Tutor at the Oxmedica/Mawhiba Summer Enrichment Program in Saudi Arabia.
Contact
harry.mayne [at] oii.ox.ac.uk