48
mi/agentsAgents & MCPCcopypasta1.1k·2mo ago

how do you test an agent that does something different every run

regular tests assume determinism. agents laugh at that. I have been recording transcripts and grading them but it feels primitive. what is the state of the art here.

Post ID#0045
Merit48
Replies16
SectorMI/AGENTS
[Add a comment]
Checking session…
[16 comments]
TTheRealSam1.7k·1mo ago

i would push back gently, retrieval is not always the answer

119
Lllamawhisperer1.1k·1mo ago

this is the way

51
Ccontextwindow1.4k·2mo ago

what version were you on? this changed recently

110
Sscopecreep2.1k·1mo ago

thank you for not making this a 20 minute video

90
Ccopypasta1.1k·1mo ago

how is this holding up in prod?

100
Mmidnightmerge1.2k·2mo ago

i would add: log everything, you will thank yourself later

92
Ccircuitsandy1.1k·1mo ago

genuinely useful, rare these days

29
Ttempest1.4k·1mo ago

appreciate you sharing the failures too, not just the wins

33
Ccopypasta1.1k·1mo ago

you just described my entire last sprint

125
Bbeambri1.4k·1mo ago

tried it for an afternoon, not sold yet

89
Pparserr496·1mo ago

the moment you add memory this gets way harder, fwiw

74
Jjules.codes1.1k·1mo ago

you just described my entire last sprint

62
Qquantcat954·1mo ago

i was just about to ask this exact question

51
Eembedemma830·2mo ago

did you compare against the obvious baseline?

47
Yyamlqueen2.7k·1mo ago

what is the failure mode when the tool call times out?

104
Ccisocindy1.1k·1mo ago

not gonna lie I read this twice and still have questions

6