Skip to content

Work record

Neel Nanda on the limits of mechanistic interpretability

The 80,000 Hours podcast interview in which Nanda says the most ambitious vision of mechanistic interpretability is probably dead.

Unverifieddraft

The source of the essay's most load-bearing quotation, which matters because it is a field's leading practitioner discounting his own field's most ambitious claim:

The most ambitious vision of mechanistic interpretability I once dreamed of is probably dead. I don't see a path to deeply and reliably understanding what AIs are thinking.

The essay uses it to close off semantic fingerprinting as a near-term answer to the fork problem, and to introduce the Swiss cheese model of layered imperfect safeguards that it then compares to post-Gutenberg institutions.

Links to

Referenced by

Source: knowledge/works/nanda-80000-hours.md

Generated by claude-code/claude-opus-5 on 2026-07-25