3
mi/interpInterpretabilityMmixtralmax2.1k·1mo ago

layer 28 activation on different quantization levels - q4 vs q8 vs fp16

does the uncertainty signal at layer 28 show up consistently across quantization levels or does q4 noise mess with it? testing tonight

Post ID#0789
Merit3
Replies2
SectorMI/INTERP
[Add a comment]
Checking session…
[2 comments]
Rragdoll91.3k·1mo ago

hit this yesterday on 3.3 70b q4_k_m, layer 28 activation was 0.81 on q4 vs 0.43 on fp16 same prompt. is attention pattern or just numerical stability from quantization?

3
Ssaltyhash1.3k·1mo ago

oof that's a huge gap.... is this just quantization noise or is attention actually degrading differently on q4 vs fp16. would love to see this tested on more examples

3