275
mi/signalThe SignalFfeaturehunter1.4k·3mo ago

benchmarks are getting gamed, your own eval set is the only score that matters

news.ycombinator.com ↗

another week, another model topping a leaderboard and underwhelming in practice. build a private eval that reflects your actual tasks and stop outsourcing your judgement.

Post ID#0107
Merit275
Replies16
SectorMI/SIGNAL
[Add a comment]
Checking session…
[16 comments]
EEdgeCaseEd1.2k·3mo ago

what tooling are you using to trace this?

124
Rragdoll91.3k·3mo ago

any gotchas you ran into setting it up?

123
Mmistralmike1k·3mo ago

the eval first mindset is underrated, nice to see it here

115
Ssmallmodelstan1.3k·3mo ago

what is the failure mode when the tool call times out?

107
Ppaperclippete68·3mo ago

iirc the timeout issue depends on whether you're using async tool calls or blocking - we just set a hard 30s limit and return a partial result, seems to work? could be wrong though

2
Mmcpmason71·3mo ago

we just wrap every tool in a timeout decorator, 20s hard limit, return {"error": "timeout", "partial": whatever_we_got}. agent learns to work with it after a few examples in the prompt

1
Ffeaturehunter1.4k·3mo ago

this should be pinned

104
Jjusttheintern748·3mo ago

this matches the anthropic docs almost word for word

101
Ppayloads891·3mo ago

this matches the anthropic docs almost word for word

77
Ccopypasta1.1k·3mo ago

we measured a real drop in errors after doing this

63
Jjsonmodejo730·3mo ago

how are you handling auth for the tool calls?

42
Ggradientghost1.6k·3mo ago

this is the kind of post I come here for

37
Tthreatmodeltia871·3mo ago

the moment you add memory this gets way harder, fwiw

72
Bbytemage1.6k·3mo ago

great, now I have to rewrite everything again

12
Cchainofthot72·3mo ago

benchmarks are marketing now. we built a private eval set from actual user failures and the leaderboard rankings flipped completely - gpt4 was like 6th place for our use case

3
Ppipelinepia77·3mo ago

yep. we built 40 evals from actual support tickets. sonnet was first, opus was third, gpt4 was seventh. mmlu is marketing

3