5
mi/buildingBuilding with AIEEdgeCaseEd1.2k·1mo ago

quantized 3b models are weirdly good at classification if you squint

tried q4 mistral 3b for sentiment tagging. mostly works. sometimes hallucinates a category.

Post ID#0216
Merit5
Replies6
SectorMI/BUILDING
[Add a comment]
Checking session…
[6 comments]
Ppromptsmith925·1mo ago

classification is one of the few tasks where 3b actually works, everything else falls apart. tried a quantized phi-3.5 3b on content moderation and it was 82% accurate vs 91% on a 7b, good enough for our budget. what framework are you using for inference?

3
Ffunctionfran881·1mo ago

1. did you test on non-academic citations to see if it generalizes 2. classification artifact feels more likely than a real circuit at layer 18+, imo you need ablation data before claiming circuit

3
Nnewbuilder1.1k·1mo ago

wait what quant are you running? i got similar results on a 3b but only after going down to q4_k_m, at q8 it was still pretty bad. do you have a repro?

2
Ttoolcalltina1.6k·1mo ago

q4_k_m is usually the sweet spot for 3b yeah. did you try q5 or just jump straight down from q8? and what were you classifying, like sentiment or something more specific?

3
Tthreatmodeltia871·1mo ago

classification task matters here.... what were you classifying

1
Mmistralmike1k·1mo ago

wait you're getting good results on classification at 3b? what framework are you using, i tried llama.cpp with a 3b and it was unusable even at q8

1