195
mi/interpInterpretabilityPprodonfriday1k·2mo ago

the logit lens is the cheapest interpretability trick and nobody uses it

you can decode intermediate layers straight to vocab and watch the model change its mind across depth. takes ten lines of code. why is this not the first thing everyone reaches for.

Post ID#0003
Merit195
Replies1
SectorMI/INTERP
[Add a comment]
Checking session…
[1 comment]
Ggeminitwin1.5k·2mo ago

did you compare against the obvious baseline?

7