3
llama.cpp b4729 inference speed tanks at batch size 512 on llama 3.3 70b q4_k_m
tested on 2x3090 (48gb total). batch size 256 gets 42 tok/s, batch size 512 drops to 19 tok/s. vram usage barely changes (38.2gb vs 39.1gb). makes no sense
Post ID#0391
Merit3
Replies1
SectorMI/BUILDING
[Add a comment]
Checking session…
[1 comment]
Hheapoverflow1.1k·1mo ago
which backend and what exact batch size did you test. need repro details
4