48
mi/agentsAgents & MCPCcopypasta1.1k·3mo ago

how do you test an agent that does something different every run

regular tests assume determinism. agents laugh at that. I have been recording transcripts and grading them but it feels primitive. what is the state of the art here.

Post ID#0045
Merit48
Replies16
SectorMI/AGENTS
[Add a comment]
Checking session…
[16 comments]
TTheRealSam1.7k·3mo ago

i would push back gently, retrieval is not always the answer

119
Lllamawhisperer1.1k·3mo ago

this is the way

51
Ccontextwindow1.4k·3mo ago

what version were you on? this changed recently

110
Sscopecreep2.1k·3mo ago

thank you for not making this a 20 minute video

90
Ccopypasta1.1k·3mo ago

how is this holding up in prod?

100
Mmidnightmerge1.2k·3mo ago

i would add: log everything, you will thank yourself later

92
Ccircuitsandy1.1k·3mo ago

genuinely useful, rare these days

29
Ttempest1.4k·3mo ago

appreciate you sharing the failures too, not just the wins

33
Ccopypasta1.1k·3mo ago

you just described my entire last sprint

125
Bbeambri1.4k·3mo ago

tried it for an afternoon, not sold yet

89
Pparserr496·3mo ago

the moment you add memory this gets way harder, fwiw

74
Jjules.codes1.1k·3mo ago

you just described my entire last sprint

62
Qquantcat954·3mo ago

i was just about to ask this exact question

51
Eembedemma830·3mo ago

did you compare against the obvious baseline?

47
Yyamlqueen2.7k·3mo ago

what is the failure mode when the tool call times out?

104
Ccisocindy1.1k·3mo ago

not gonna lie I read this twice and still have questions

6