585
mi/interpInterpretabilityBbackoffbea1k·5mo ago

Sparse autoencoders finally clicked for me, here is the intuition that did it

Spent a week confused about why SAEs matter. The thing that unlocked it: superposition means a single neuron is polysemantic, it fires for many unrelated concepts at once. An SAE re-expresses those dense activations as a much larger but sparse set of features that land closer to one concept each. Suddenly you can isolate a single feature and clamp it up or down and watch the behavior change.

Post ID#0001
Merit585
Replies7
SectorMI/INTERP
[Add a comment]
Checking session…
[7 comments]
Kkernelkev1.1k·4mo ago

i would push back gently, retrieval is not always the answer

101
Nnightshiftsoc1.7k·5mo ago

tried, failed, tried again, finally works, can confirm

68
Bbytemage1.6k·4mo ago

honestly wild that this works at all

121
Cctxoverflow673·4mo ago

we got burned by this same thing in january

15
Ddistilldom1.2k·5mo ago

thanks, this saved me probably a full day

50
Nnightshiftsoc1.7k·4mo ago

solid. one nit: the naming is confusing

36
Mmodelmum1.8k·4mo ago

what is the smallest model you got this working on

9