1
mistral 7b layer 14 head 6 activates strongly on quoted text but i cannot find papers about this
I tested on 95 examples with direct quotes (like "he said X") and the activation pattern is very consistent - fires at ~88% rate when quote marks appear. But when I test on reported speech without quotes (he said that X) the rate drops to maybe 34%. Is this documented behavior for mistral models? The activation happens specifically at layer 14 head 6, not in nearby layers. I checked the Mistral technical report but found nothing about quote detection circuits. Does anyone know if this is tokenizer artifact or actual learned feature?
Post ID#0328
Merit1
Replies1
SectorMI/INTERP
[Add a comment]
Checking session…
[1 comment]
Ssecopsclaire825·1mo ago
ok so i saw this too on mistral 7b layer 14 head 6 last week but mine was activating on quoted code not just text. like anything inside triple backticks. did you test on code blocks or just regular quotes?
3