historical research · archived

Interpretability Frontier

Can Mechanistic Interpretability Substitute for Structural Diagnosis?

Does component-level interpretability recover composition-level semantics?

The controlled experiment remains reported. The paper family is archived and is not a current priority.

Abstract

Tests whether mechanistic interpretability (SAE, probing, circuit tracing) can substitute for structural diagnosis on cyclic compositional failure.

Record
Status
archivedEmpirically demonstratedhistorical · v1 · as of 2026-07-22
Corrections

Archived to enforce the no-fourth-family rule during the reset.

Falsification

not-applicable

External review

not requested

Program relations

Historical decision · 2026-07-14

Status decision

Why it was pursued
The experiment tested whether component-level interpretability recovers non-local semantics.
Evidence change
The scoped experiment did not justify treating interpretability as the program capstone, and the reset forbids a fourth active family.
No longer claimed
  • Current program capstone
  • General boundary for all interpretability systems
Retained results
  • Scoped reported experiment
  • Representation-sensitive design lesson