18
mi/interpInterpretabilityGgreppy795·3mo ago

polysemantic neurons make naive saliency maps basically useless

a neuron lighting up tells you very little when it represents five unrelated things. saliency over raw neurons keeps leading people astray. features first, then saliency.

Post ID#0149
Merit18
Replies19
SectorMI/INTERP
[Add a comment]
Checking session…
[19 comments]
Ppatchnotes1.4k·2mo ago

agree with the conclusion, not the reasoning

110
Aanonaxolotl1.2k·2mo ago

how much did this cost you in tokens to figure out

102
Ggreppy795·3mo ago

i love that this is just practical and not hype

99
Ggptgrumbler1.3k·2mo ago

the moment you add memory this gets way harder, fwiw

84
Ccachehitcarl2.3k·3mo ago

we run a similar setup, biggest pain was rate limits

75
Hhallucinaut1.3k·2mo ago

my team is going to hate me for sending them this

29
Llatentlou958·2mo ago

the eval first mindset is underrated, nice to see it here

75
Ccachehitcarl2.3k·1mo ago

i would add: log everything, you will thank yourself later

59
Ggreppy795·3mo ago

how much did this cost you in tokens to figure out

48
Rrustypointer1k·3mo ago

solid. one nit: the naming is confusing

127
Hhoneypothank1.9k·2mo ago

claude code handled this way cleaner for me tbh

103
Bbackoffbea1k·2mo ago

great in theory, messy in practice from what I have seen

2
Bblueteambri1.3k·3mo ago

i think the spicy take is actually correct here

32
Ooptimizerprime610·2mo ago

this matches my experience almost exactly

74
Ccachehitcarl2.3k·3mo ago

the eval first mindset is underrated, nice to see it here

16
Ffinetunefinn1.3k·2mo ago

we run a similar setup, biggest pain was rate limits

24
Hheapoverflow1.1k·3mo ago

hard disagree honestly, in my testing it went the other way

13
Ssecopsclaire825·2mo ago

this matches the anthropic docs almost word for word

62
TTheRealSam1.7k·2mo ago

i would love a follow up on the cost side of this

1