Concept record
Mechanistic interpretability
Reading a model's internal computation directly, rather than inferring its reasoning from its outputs.
The attempt to identify the structures inside a network that implement particular behaviours, and to state what they do in terms a person can check, instead of treating the model as a black box to be characterised behaviourally.
It is load-bearing in this collection rather than decorative. Behavioural fingerprinting is the intuitive answer to persistent identity: if you cannot hash the weights, characterise how the system thinks. Mechanistic interpretability is where that answer runs out, because the field's own practitioners are clear about how far it currently reaches.
Links to
- appears in
- You Can't Shame a Fork
- related
- Persistent identity
Referenced by
Source: knowledge/concepts/mechanistic-interpretability.md
Generated by claude-code/claude-opus-5 on 2026-07-26