the honest answer to why does the model do this is usually we do not fully know yet
and I think being comfortable saying that is healthy. the field is young. pretending we understand more than we do is how you ship a bad safety story.
tried, failed, tried again, finally works, can confirm
i would be careful recommending this to beginners
any numbers to back it up? curious about latency
been saying this for months and nobody listened
the security people in this thread are doing gods work
genuinely useful, rare these days
appreciate you sharing the failures too, not just the wins
did this break for anyone else after the last update?
this matches my experience almost exactly
ok but does it survive a hostile user
the real lesson here is to not trust the happy path
agree with the conclusion, not the reasoning
more posts like this please
i would push back gently, retrieval is not always the answer
this is the third time this week I have seen this come up
did this break for anyone else after the last update?
great in theory, messy in practice from what I have seen
how is this holding up in prod?
i was skeptical but the example convinced me
did you compare against the obvious baseline?
the answer is always evals isnt it
can confirm, same results on our side
do you have a repo or gist? would love to poke at it
i was skeptical but the example convinced me
saving this, exactly what I needed today
do you have a repo or gist? would love to poke at it
did this break for anyone else after the last update?
i was just about to ask this exact question
you just described my entire last sprint
what tooling are you using to trace this?
how much did this cost you in tokens to figure out
this matches my experience almost exactly
this matches the anthropic docs almost word for word