bibkey: “bricken2023monosemanticity” authors: “Trenton Bricken; Adly Templeton; Joshua Batson; Brian Chen; Adam Jermyn; Tom Conerly; Nicholas L. Turner; Cem Anil; Carson Denison; Amanda Askell; Robert Lasenby; Yifan Wu; Shauna Kravec; Nicholas Schiefer; Tim Maxwell; Nicholas Joseph; Alex Tamkin; Karina Nguyen; Brayden McLean; Josiah E. Burke; Tristan Hume; Shan Carter; Tom Henighan; Chris Olah” year: 2023 title: “Towards Monosemanticity: Decomposing Language Models With Dictionary Learning” doi: null claim: “Sparse autoencoders trained on transformer activations yield interpretable features; larger dictionaries split features into more specific ones.” strata_touched: [] license: “citation-only” triage: “anchor” url: “https://transformer-circuits.pub/2023/monosemantic-features/index.html”
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Verified locator
Transformer Circuits Thread, October 2023. Cited as background for sparse autoencoders and feature splitting; no theorem from it is used in the fiber-law volume.
Declared identifiers: https://transformer-circuits.pub/2023/monosemantic-features/index.html.