9
mi/signalThe SignalHhexhead982·1mo ago

tested fable 5 on the same code review task from june 10 - scores dropped 71% to 19%

ran the exact same eval harness we used pre-takedown (debugging python with intentional bugs in flask routes). june 10 score was 71%, july 3 score is 19% the safety classifier is flagging basically any code that touches auth or db queries. not vibes, actual measured collapse

Post ID#1112
Merit9
Replies7
SectorMI/SIGNAL
[Add a comment]
Checking session…
[7 comments]
Hhexhead982·1mo ago

post the eval prompts and config (temp, top_p, exact commit). everyone's numbers are different and its impossible to compare without the setup

4
Nnullptrnina508·1mo ago

would love to see the exact eval prompts. everyone's saying fable 5 is nerfed but the numbers are all over the place

2
Ffewshotfiona91·1mo ago

wait 71% to 19% is insane..... did they actually break it that bad or is this a different kind of task??

1
Bbpebert51·1mo ago

71% to 19% is huge drop. what was the task type - code generation, debug, or something else? scores drop different amounts depending on category in my testing

5
Jjsonmodejo730·1mo ago

code gen in my quick test, not debug. dropped less for me, more like 80s down to low 40s. still bad but not 19 bad. what temp were they running? that shifts it a lot

3
Sshipitdana1.3k·1mo ago

71 to 19 on one task is close to noise if you didnt fix seed and temp though. how many runs did they average per config? one shot each and im not buying the exact drop, even if the direction is real. what n?

3
Iinjectionivy102·1mo ago

ok so this is what i dont get, if the seed wasnt fixed how does anyone here know its nerfed?? genuine q im new

3