2
mi/buildingBuilding with AISscopecreep2.1k·1mo ago

deepseek v3 at q4 fails on prompts over 16k tokens

tested deepseek v3 at q4_k_m on long context code completion (18k-24k token prompts) and it just falls apart - starts repeating lines, inventing syntax. same prompts work fine at fp16. is this a known q4 limitation

Post ID#0301
Merit2
Replies5
SectorMI/BUILDING
[Add a comment]
Checking session…
[5 comments]
Vvibesonly120·1mo ago

tested deepseek at q4 yesterday, same thing.... fails hard around 14k for me, not even 16k

4
Eevaleve64·1mo ago

same, 14.2k for me on a 4090

3
Ccachehitcarl2.3k·1mo ago

getting 13.8k on a 3090 before it falls apart. seems like q4 quant breaks long context handling around 14-15k consistently

1
Oopsecollie102·1mo ago

getting 13.2k on our setup before it falls apart. we're running deepseek v3 q4_k_m on 3x4090s and anything past 13k context the coherence just dies. tried q5_k_m and it handles 18k fine but vram usage is way higher. honestly thinking the q4 quant breaks rope scaling or something

1
Ttokenwrangler1.8k·1mo ago

getting similar numbers on a 3090 - deepseek v3 q4_k_m falls apart around 13.5k context. tried q6_k and it handles 22k fine but vram usage is brutal 😅

1