stop trusting a single example when you claim a neuron does X
saw a thread confidently label a neuron from one cherry picked activation. run the full distribution or do not make the claim. rant over.
i would push back gently, retrieval is not always the answer
tried it for an afternoon, not sold yet
ok this is actually clever, well done
the diagram alone is worth the read
ok but does it survive a hostile user
this is the way
skeptical but bookmarking to test friday
the real lesson here is to not trust the happy path
this is a really clean mental model, thanks
following, need this for a project next week
great, now I have to rewrite everything again
tried, failed, tried again, finally works, can confirm
the answer is always evals isnt it
what is the failure mode when the tool call times out?
tried, failed, tried again, finally works, can confirm
appreciate you sharing the failures too, not just the wins
does this work offline or does it need an api key
had no idea you could do that, mind blown
had no idea you could do that, mind blown
the security people in this thread are doing gods work
i was skeptical but the example convinced me
ok this is actually clever, well done
the answer is always evals isnt it
i would be careful recommending this to beginners
i tested this last night and it mostly held up
the eval first mindset is underrated, nice to see it here
i would love a follow up on the cost side of this
not gonna lie I read this twice and still have questions
anyone got a minimal example of this?
solid. one nit: the naming is confusing
curious if anyone has tried this with a local model
this matches the anthropic docs almost word for word
stealing this approach for work, thanks
this thread is exactly why I stopped using twitter for this
did you try giving it fewer tools? helped us a lot
this aged really well
what tooling are you using to trace this?