Interpretability Frontier
Can Mechanistic Interpretability Substitute for Structural Diagnosis?
Does component-level interpretability recover composition-level semantics?
The controlled experiment remains reported. The paper family is archived and is not a current priority.
Abstract
Tests whether mechanistic interpretability (SAE, probing, circuit tracing) can substitute for structural diagnosis on cyclic compositional failure.
Record
Status
archivedEmpirically demonstratedhistorical · v1 · as of 2026-07-22
Corrections
Archived to enforce the no-fourth-family rule during the reset.
Falsification
not-applicable
External review
not requested
Program relations
- supersedes Compositional Accountability
Historical decision · 2026-07-14
Status decision
- Why it was pursued
- The experiment tested whether component-level interpretability recovers non-local semantics.
- Evidence change
- The scoped experiment did not justify treating interpretability as the program capstone, and the reset forbids a fourth active family.
- No longer claimed
- Current program capstone
- General boundary for all interpretability systems
- Retained results
- Scoped reported experiment
- Representation-sensitive design lesson
SuccessorsCompositional Accountability