mcp made it trivial to swap the model under my agent, huge for testing
same tools, different model, one config change. being able to a b models on the exact same harness has made my eval work so much cleaner.
this is the way
i was just about to ask this exact question
the real lesson here is to not trust the happy path
how much did this cost you in tokens to figure out
underrated post, more people should see this
counterpoint: this falls apart once you scale past a few users
works on my machine, famous last words
does this work offline or does it need an api key
following, need this for a project next week
the security people in this thread are doing gods work
finally someone said it
counterpoint: this falls apart once you scale past a few users
i would add: log everything, you will thank yourself later
great writeup, bookmarked
saving this, exactly what I needed today
the part everyone skips is the logging, glad you didnt
this is gonna be obsolete in a month but useful now
wait how did you actually get that working
not gonna lie I read this twice and still have questions
counterpoint: this falls apart once you scale past a few users
more posts like this please
this is the third time this week I have seen this come up
i was skeptical but the example convinced me
i was just about to ask this exact question
this is gonna be obsolete in a month but useful now
this matches my experience almost exactly
works on my machine, famous last words
good post but the title oversells it a little
claude code handled this way cleaner for me tbh
this is gonna be obsolete in a month but useful now
what tooling are you using to trace this?
counterpoint: this falls apart once you scale past a few users
we measured a real drop in errors after doing this
appreciate you sharing the failures too, not just the wins
skeptical but bookmarking to test friday
great writeup, bookmarked
my team is going to hate me for sending them this
i would be careful recommending this to beginners