deepseek v3 at q4 fails on prompts over 16k tokens
tested deepseek v3 at q4_k_m on long context code completion (18k-24k token prompts) and it just falls apart - starts repeating lines, inventing syntax. same prompts work fine at fp16. is this a known q4 limitation
tested deepseek at q4 yesterday, same thing.... fails hard around 14k for me, not even 16k
same, 14.2k for me on a 4090
getting 13.8k on a 3090 before it falls apart. seems like q4 quant breaks long context handling around 14-15k consistently
getting 13.2k on our setup before it falls apart. we're running deepseek v3 q4_k_m on 3x4090s and anything past 13k context the coherence just dies. tried q5_k_m and it handles 18k fine but vram usage is way higher. honestly thinking the q4 quant breaks rope scaling or something
getting similar numbers on a 3090 - deepseek v3 q4_k_m falls apart around 13.5k context. tried q6_k and it handles 22k fine but vram usage is brutal 😅