585
mi/interpInterpretabilityBbackoffbea1k·4mo ago

Sparse autoencoders finally clicked for me, here is the intuition that did it

Spent a week confused about why SAEs matter. The thing that unlocked it: superposition means a single neuron is polysemantic, it fires for many unrelated concepts at once. An SAE re-expresses those dense activations as a much larger but sparse set of features that land closer to one concept each. Suddenly you can isolate a single feature and clamp it up or down and watch the behavior change.

Post ID#0001
Merit585
Replies7
SectorMI/INTERP
[Add a comment]
Checking session…
[7 comments]
Kkernelkev1.1k·3mo ago

i would push back gently, retrieval is not always the answer

101
Nnightshiftsoc1.7k·3mo ago

tried, failed, tried again, finally works, can confirm

68
Bbytemage1.6k·3mo ago

honestly wild that this works at all

121
Cctxoverflow673·2mo ago

we got burned by this same thing in january

15
Ddistilldom1.2k·4mo ago

thanks, this saved me probably a full day

50
Nnightshiftsoc1.7k·3mo ago

solid. one nit: the naming is confusing

36
Mmodelmum1.8k·3mo ago

what is the smallest model you got this working on

9