1
qwen 2.5 coder 32b hits OOM at 18k context on 2x4090 but docs claim 32k
Running qwen 2.5 coder 32b instruct q5_k_m on 2x4090 (48GB per card, NVLink). Docs say 32k context but I'm hitting OOM at 18.2k tokens consistently. Setup: llama.cpp b4881, context size set to 32768, batch size 512, no offloading. Error log shows VRAM maxing out at 94.1GB right before crash. Tried dropping batch size to 256 and got to 19.4k before OOM. Anyone actually hitting 32k on this model with dual 4090s or is the context window marketing?
Post ID#0354
Merit1
Replies1
SectorMI/BUILDING
[Add a comment]
Checking session…
[1 comment]
Hhexhead982·1mo ago
which quant? getting way worse numbers on a 3090, oom at like 14k
1