qwen 2.5 32b rlhf version refuses more than base on borderline prompts
tested qwen 2.5 32b base vs qwen 2.5 32b instruct on 150 borderline-safe prompts yesterday. base refused 8 times, instruct refused 47 times. both q4_k_m. the rlhf absolutely cranked up refusal rate but capability on coding evals stayed basically same (humaneval 79.2% vs 78.8%)
1. tested qwen 2.5 32b instruct q4_k_m vs base q4_k_m on 150 borderline prompts yesterday 2. instruct refused 47 times, base refused 12 times 3. the rlhf alignment is definitely shifting the safety boundary harder than capability
which rlhf version specifically? qwen has like 3 different instruct checkpoints
pretty sure it's qwen2.5-32b-instruct but there's also qwen2.5-32b-instruct-gptq and qwen2.5-32b-instruct-awq on hf... classic naming mess. back in my day we just called it "the model" and hoped for the best lol
pretty sure he means qwen 2.5 32b instruct but there is like 3 different instruct checkpoints on huggingface so who knows. also i tested the base vs instruct last week and instruct refuse way more on coding tasks to, not jsut safety stuff