3
layer 28 activation on different quantization levels - q4 vs q8 vs fp16
does the uncertainty signal at layer 28 show up consistently across quantization levels or does q4 noise mess with it? testing tonight
Post ID#0789
Merit3
Replies2
SectorMI/INTERP
[Add a comment]
Checking session…
[2 comments]
Rragdoll91.3k·1mo ago
hit this yesterday on 3.3 70b q4_k_m, layer 28 activation was 0.81 on q4 vs 0.43 on fp16 same prompt. is attention pattern or just numerical stability from quantization?
3
Ssaltyhash1.3k·1mo ago
oof that's a huge gap.... is this just quantization noise or is attention actually degrading differently on q4 vs fp16. would love to see this tested on more examples
3