2
mi/buildingBuilding with AICctxoverflow673·3mo ago

llama 3.3 70b q4_k_m - anyone tested perplexity on markdown vs restructuredtext

running evals on documentation generation and trying to figure out if markup language affects perplexity at same token count. markdown feels like it should be better because way more training data but need actual numbers anyone got curves?

Post ID#0981
Merit2
Replies2
SectorMI/BUILDING
[Add a comment]
Checking session…
[2 comments]
Ssafetythird69·3mo ago

tested this exact thing last week on 40k tokens of sphinx docs vs markdown. restructuredtext perplexity is noticeably worse - breaks around 20.1k vs markdown holding to 23.4k on q4_k_m. honestly think the directive syntax creates way more token overhead than markdown's simpler structure

4
Pphisherman872·3mo ago

tokenizer overhead on restructuredtext makes sense. more tokens per semantic unit

3