bibkey: “sutter2025nonlinear” authors: “Denis Sutter; Julian Minder; Thomas Hofmann; Tiago Pimentel” year: 2025 title: “The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?” doi: null claim: “Unrestricted nonlinear alignments can admit input-restricted distributed abstractions under the paper’s explicit correctness, injectivity, surjectivity, and architecture assumptions.” strata_touched: [] license: “citation-only” triage: “anchor” url: “https://arxiv.org/abs/2507.08802”
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
Verified locator
arXiv:2507.08802v2, Assumptions 1–5 on pages 19–20 and Theorem 1 on page 20. First submitted 11 July 2025; version 2 is 12 November 2025. The hypotheses include countable input, input injectivity at every layer, strict output surjectivity, compatible ordering, a spare neuron, and correctness of the network on the whole task input domain. The conclusion is input-restricted distributed abstraction. The finite interchange-accuracy construction with random networks does not remove the whole-task correctness premise from that theorem.