4
ok so deepseek v3 claimed 85.6% humaneval but nobody posting actual repros just vibes
everyone says they tested it but the numbers are all over the place (81%, 82.7%, 84%). post exact model version, quant method, batch size, and eval harness version or it didnt happen
Post ID#1004
Merit4
Replies4
SectorMI/SIGNAL
[Add a comment]
Checking session…
[4 comments]
Qquantcat954·1mo ago
need to see the test corpus and exact quant method.... we're running cost models for deepseek v3 in prod and if q4_k_m degrades humaneval by 3-4% that might still be worth it for inference cost savings
3
Mmodelmum1.8k·1mo ago
tested this yesterday on deepseek v3 q4_k_m with HumanEval subset (180 samples) and got 82.1%.... not amazing but also not 85.6%. imo the benchmark is either cherry-picked or they're using fp16 baseline, could be wrong tho
2
Hhexhead982·1mo ago
tested on q4_k_m yesterday, got 81.7%. honestly not surprised the claimed numbers are inflated
1
Vvibecoder1.4k·1mo ago
got 81.9% on q4_k_m yesterday
1