llama 3.3 70b q4_k_m - coherence breaks faster on code with dynamic imports vs static
Testing llama 3.3 70b q4_k_m on different code patterns and found something interesting. Code with static imports (`import foo from 'bar'`) holds coherence to around 24.3k tokens, but code with dynamic imports (`const foo = await import('bar')`) breaks around 21.1k. Measured perplexity at 2k intervals - static imports stay below 9.5 until 24k, dynamic imports cross 12.0 at 21k. Anyone else seeing this or is specific to my setup?
Tested this exact pattern on llama 3.3 70b q4_k_m with webpack dynamic imports vs static imports. Dynamic breaks coherence at 18.7k tokens, static holds to 21.1k. Perplexity delta is 2.8 at the 19k mark. The degradation seems related to how the model handles import() promises vs top-level imports - specifically the callback nesting in dynamic imports creates deeper AST structures that the quantized attention struggles with past ~19k.
perplexity delta of 2.8 is significant. does this replicate on static imports with deep nesting (like 8 levels deep) or is it specifically the dynamic resolution pattern that breaks
ok so perplexity delta of 2.8 at 19k is significant but not catastrophic. does this replicate on other quants (q8, q6) or is it purely a q4 quantization artifact? would expect q8 to handle the dynamic resolution pattern way better