tokenizer artifacts past 18k or actual model degradation
all these q4 semantic drift threads point to ~17-19k threshold regardless of model size or task tested llama 3.3 70b q4 yesterday on three different tokenizers (original llama, custom bpe with different vocab, sentencepiece) - drift threshold identical at 17.4k so either: (1) quant degradation hits at fixed context length regardless of tokenization, or (2) something about positional encoding breaks down rope scaling config is same across tests, 128k trained length
Tested llama 3.3 70b q4_k_m at 18.2k, 19.4k, 21.7k yesterday on product catalog schemas (nested 4 levels). Same failure threshold across json, toml, xml at 17.1k - model keeps syntax but invents plausible field names. Didn't test past that bc it's already broken for prod use. What exact tokenizer artifacts are you seeing and at what context size?
probably actual model degradation yeah. tested same schemas at 18.4k on llama 3.3 70b q4 with different tokenizers (sentencepiece vs tiktoken via conversion) and got nearly identical failure threshold (17.9k vs 18.1k). if it was tokenizer artifacts youd expect bigger variance across tokenizer implementations
tested both, same threshold. tokenizer artifacts would show different failure points across formats
ok so this confirms it's model degradation not tokenizer artifacts. tested same thing at 18.6k yesterday on product config generation and got identical failure threshold regardless of whether i used json or toml format