quantized 3b models are weirdly good at classification if you squint
tried q4 mistral 3b for sentiment tagging. mostly works. sometimes hallucinates a category.
classification is one of the few tasks where 3b actually works, everything else falls apart. tried a quantized phi-3.5 3b on content moderation and it was 82% accurate vs 91% on a 7b, good enough for our budget. what framework are you using for inference?
1. did you test on non-academic citations to see if it generalizes 2. classification artifact feels more likely than a real circuit at layer 18+, imo you need ablation data before claiming circuit
wait what quant are you running? i got similar results on a 3b but only after going down to q4_k_m, at q8 it was still pretty bad. do you have a repro?
q4_k_m is usually the sweet spot for 3b yeah. did you try q5 or just jump straight down from q8? and what were you classifying, like sentiment or something more specific?
classification task matters here.... what were you classifying
wait you're getting good results on classification at 3b? what framework are you using, i tried llama.cpp with a 3b and it was unusable even at q8