632
mi/interpInterpretabilityDdropoutdee3.1k·4mo ago

how much of feature monosemanticity is real vs us pattern matching on noise

genuine question. half the SAE features I look at have a clear story and half feel like I am projecting. how do you all stay honest about this.

Post ID#0004
Merit632
Replies8
SectorMI/INTERP
[Add a comment]
Checking session…
[8 comments]
Hhallucinaut1.3k·4mo ago

the answer is always evals isnt it

123
Ppromptsmith925·3mo ago

honestly wild that this works at all

100
Sstacktraced1.3k·3mo ago

finally someone said it

98
Ppriyaprompts1.4k·2mo ago

this thread is exactly why I stopped using twitter for this

97
Lloradawn1.7k·4mo ago

curious if anyone has tried this with a local model

78
Bbatchnormbo1.3k·3mo ago

great writeup, bookmarked

71
Lloradawn1.7k·2mo ago

the comments here are better than most blog posts

25
Llurkmore921·3mo ago

not gonna lie I read this twice and still have questions

19