Against Almost Every Theory of Impact of Interpretability
Against Almost Every Theory of Impact of Interpretability
LessWrong post critiquing most theories of impact for interpretability research in AI alignment. Referenced in the generalist reading list primarily for Richard Ngo's comment, which extracts a generalizable strategic principle about how to reason about forward-chaining work.
The original post (domain-specific, 602 lines)
Claims defended:
- The overall theory of impact for interpretability is poor
- Auditing deception with interpretability is out of reach (deceptive plans won't be in interpretable circuits)
- Proposed end-state visions (enumerative safety, reverse engineering, Olah's dream, retargeting the search) are not clearly achievable
- Best available theory of impact for interpretability: outreach; preventive measures against deception are more workable
- Closing argument: "don't build powerful dangerous AIs; prevention is better than cure"
The structure is a whack-a-mole critique — for each proposed mechanism, show why it doesn't deliver the promised safety outcome.
Richard Ngo's comment — the generalizable principle
Strong disagree from Ngo, on grounds of type of reasoning:
"This seems like very much the wrong type of reasoning to do about novel scientific research. Big breakthroughs open up possibilities that are very hard to imagine before those breakthroughs — imagine trying to describe the useful applications of electricity before anyone knew what it was or how it worked; or imagine Galileo trying to justify the practical use of studying astronomy."
"Interpretability seems like our clear best bet for developing a more principled understanding of how deep learning works; this by itself is sufficient to recommend it."
"I think this post is an example of a fairly common phenomenon where alignment people are too focused on backchaining from desired end states, and not focused enough on forward-chaining to find the avenues of investigation which actually allow useful feedback and give us a hook into understanding the world better. By contrast, most ML researchers are too focused on the latter."
The general principle: Winning whack-a-mole — showing each proposed theory of impact fails — does not conclusively discredit a forward-chaining approach. The absence of a pre-specified theory of impact does not mean an avenue is unworthy. Big breakthroughs routinely open possibilities that were unimaginable in advance. The right justification for a forward-chaining avenue is that it provides real feedback, hooks into understanding, and represents our best bet for principled knowledge in a domain.
The pathology Ngo identifies — excessive backchaining — generates critiques that feel conclusive but aren't, because they assume the world must work through already-imagined mechanisms.
Related: forward-chaining-vs-backchaining a-reading-list-for-generalists-dylan-bowman