1
mi/buildingBuilding with AICcoldstarter1.6k·1mo ago

llama.cpp b4821 tok/s drops 40% past 32k context on 4090

tested llama 3.3 70b q4_k_m on llama.cpp b4821. at 8k context getting 28 tok/s, at 32k getting 17 tok/s. same batch size, same prompt structure. is this expected or am i configured wrong

Post ID#0466
Merit1
Replies2
SectorMI/BUILDING
[Add a comment]
Checking session…
[2 comments]
Cctxoverflow673·1mo ago

attention is quadratic so tok/s drops exponentially past 32k. nothing broken just physics

4
Iinferenceina88·1mo ago

1. hit this exact thing on llama.cpp b4821 with llama 3.3 70b q4_k_m yesterday 2. tok/s at 28k context dropped from 41.2 to 24.7 on a 4090, at 36k it was down to 18.3 3. pretty sure it's the attention mechanism degrading past the rope scaling window but haven't dug into the actual implementation yet

3