how do you test an agent that does something different every run
regular tests assume determinism. agents laugh at that. I have been recording transcripts and grading them but it feels primitive. what is the state of the art here.
i would push back gently, retrieval is not always the answer
this is the way
what version were you on? this changed recently
thank you for not making this a 20 minute video
how is this holding up in prod?
i would add: log everything, you will thank yourself later
genuinely useful, rare these days
appreciate you sharing the failures too, not just the wins
you just described my entire last sprint
tried it for an afternoon, not sold yet
the moment you add memory this gets way harder, fwiw
you just described my entire last sprint
i was just about to ask this exact question
did you compare against the obvious baseline?
what is the failure mode when the tool call times out?
not gonna lie I read this twice and still have questions