lora on qwen 2.5 7b collapsed after 400 steps, loss went from 0.8 to 4.2 overnight
training lora on qwen 2.5 7b for code summarization, everything was fine until step 380 then the loss just exploded to 4.2 and the model started outputting garbage. learning rate was 1e-4, rank 16, dataset is 2000 examples of python functions with docstrings. checkpoint from step 300 still works fine. no idea what happened, anyone seen this before??
lr too high probably
lr was 2e-4, dropped to 5e-5 around step 200 but loss still exploded. what lr schedule do you use for qwen loras?
lr too high imo. we hit this exact thing on qwen 2.5 7b lora for code - loss spiked around step 500 and never recovered. dropped lr from 2e-4 to 8e-5 with cosine decay and it trained clean all the way to 2000 steps. also check if you're targeting too many layers - we had to drop from all projection layers to just q/v to stop the collapse. could be wrong but lora rank 64 on a 7b might also be too aggressive
ok so what lr and rank did you end up using that worked? hitting this exact thing on a qwen lora right now and loss is spiking around step 600