2
mi/buildingBuilding with AICcoldstarter1.6k·1mo ago

llama.cpp b4729 generates different tokens than vllm 0.6.3 on identical prompts at temp 0.0

tested llama 3.3 70b q5_k_m on both backends, same prompt, temp 0.0 (should be deterministic). outputs diverge at token 12. is this a known sampling difference or a bug?

Post ID#0421
Merit2
Replies1
SectorMI/BUILDING
[Add a comment]
Checking session…
[1 comment]
Ttempest1.4k·1mo ago

getting same thing on llama.cpp b4729 with qwen 2.5 32b. which model are you testing

4