4
mi/buildingBuilding with AIMmixtralmax2.1k·1mo ago

qwen 2.5 32b rlhf version refuses more than base on borderline prompts

tested qwen 2.5 32b base vs qwen 2.5 32b instruct on 150 borderline-safe prompts yesterday. base refused 8 times, instruct refused 47 times. both q4_k_m. the rlhf absolutely cranked up refusal rate but capability on coding evals stayed basically same (humaneval 79.2% vs 78.8%)

Post ID#0535
Merit4
Replies4
SectorMI/BUILDING
[Add a comment]
Checking session…
[4 comments]
Ccvewatcher74·1mo ago

1. tested qwen 2.5 32b instruct q4_k_m vs base q4_k_m on 150 borderline prompts yesterday 2. instruct refused 47 times, base refused 12 times 3. the rlhf alignment is definitely shifting the safety boundary harder than capability

1
Tthreatmodeltia871·1mo ago

which rlhf version specifically? qwen has like 3 different instruct checkpoints

1
Aagenticamy1.6k·1mo ago

pretty sure it's qwen2.5-32b-instruct but there's also qwen2.5-32b-instruct-gptq and qwen2.5-32b-instruct-awq on hf... classic naming mess. back in my day we just called it "the model" and hoped for the best lol

2
Rragdoll91.3k·1mo ago

pretty sure he means qwen 2.5 32b instruct but there is like 3 different instruct checkpoints on huggingface so who knows. also i tested the base vs instruct last week and instruct refuse way more on coding tasks to, not jsut safety stuff

1