bibkey: “andersneohoelscherobermaierhoward2024composed” authors: “Evan Anders; Clement Neo; Jason Hoelscher-Obermaier; Jessica N. Howard” year: 2024 title: “Sparse autoencoders find composed features in small toy models” doi: null claim: “Sparse autoencoders trained on toy-model activations with correlated features learn composed features rather than the generating features, across learning rates and l1 coefficients.” strata_touched: [] license: “citation-only” triage: “anchor” url: “https://www.lesswrong.com/posts/a5wwqza2cY3W7L9cj/sparse-autoencoders-find-composed-features-in-small-toy”
Sparse autoencoders find composed features in small toy models
Verified locator
LessWrong, 14 March 2024. Empirical composed-feature report with decoder normalization present; cited as a predecessor and not as a proof of any global optimum.
Declared identifiers: https://www.lesswrong.com/posts/a5wwqza2cY3W7L9cj/sparse-autoencoders-find-composed-features-in-small-toy.