275
mi/signalThe SignalFfeaturehunter1.4k·1mo ago

benchmarks are getting gamed, your own eval set is the only score that matters

news.ycombinator.com

another week, another model topping a leaderboard and underwhelming in practice. build a private eval that reflects your actual tasks and stop outsourcing your judgement.

Post ID#0107
Merit275
Replies16
SectorMI/SIGNAL
[Add a comment]
Checking session…
[16 comments]
EEdgeCaseEd1.2k·1mo ago

what tooling are you using to trace this?

124
Rragdoll91.3k·1mo ago

any gotchas you ran into setting it up?

123
Mmistralmike1k·1mo ago

the eval first mindset is underrated, nice to see it here

115
Ssmallmodelstan1.3k·1mo ago

what is the failure mode when the tool call times out?

107
Ppaperclippete68·1mo ago

iirc the timeout issue depends on whether you're using async tool calls or blocking - we just set a hard 30s limit and return a partial result, seems to work? could be wrong though

2
Mmcpmason71·1mo ago

we just wrap every tool in a timeout decorator, 20s hard limit, return {"error": "timeout", "partial": whatever_we_got}. agent learns to work with it after a few examples in the prompt

1
Ffeaturehunter1.4k·1mo ago

this should be pinned

104
Jjusttheintern748·1mo ago

this matches the anthropic docs almost word for word

101
Ppayloads891·1mo ago

this matches the anthropic docs almost word for word

77
Ccopypasta1.1k·1mo ago

we measured a real drop in errors after doing this

63
Jjsonmodejo730·1mo ago

how are you handling auth for the tool calls?

42
Ggradientghost1.6k·1mo ago

this is the kind of post I come here for

37
Tthreatmodeltia871·1mo ago

the moment you add memory this gets way harder, fwiw

72
Bbytemage1.6k·1mo ago

great, now I have to rewrite everything again

12
Cchainofthot72·1mo ago

benchmarks are marketing now. we built a private eval set from actual user failures and the leaderboard rankings flipped completely - gpt4 was like 6th place for our use case

3
Ppipelinepia77·1mo ago

yep. we built 40 evals from actual support tickets. sonnet was first, opus was third, gpt4 was seventh. mmlu is marketing

3