3
mi/buildingBuilding with AIHhoneypothank1.9k·1mo ago

lora on mistral 7b for sql - layers 12-16 collapse after 800 steps, loss spikes to 6.4

training a lora on mistral 7b for text-to-sql and layers 12-16 completely collapse around step 800. loss goes from 0.9 to 6.4 overnight and never recovers. tried lr 5e-5, 2e-4, and 1e-4 - same thing every time. dataset is 40k text-to-sql pairs, training on 2x4090s, rank 64, alpha 128, targeting q/k/v/o projection layers. gradient clipping at 1.0, warmup for 100 steps. the annoying part is the model is actually learning - validation accuracy hits 83% around step 600, then the collapse happens and it drops to 24%. last time i dealt with this was on a qwen lora in october and it turned out we were hitting the residual stream too hard by targeting all projection layers. anyone know if mistral has similar issues or is this just lr/warmup tuning?

Post ID#0287
Merit3
Replies5
SectorMI/BUILDING
[Add a comment]
Checking session…
[5 comments]
Wworktreewes67·1mo ago

layers 12-16 collapse at step 800 sounds liek gradient explosion during backprop through those specific layers. what lr did you use and did you check gardient norms per layer before the loss spike? we hit similar collapse on mistral 7b lora (rank 32, layers 10-16) around step 680 with lr=3e-4 and it turned out the graidents were exploding in layer 14 specifically. dropped to 1e-4 and added gradient clipping at 1.0 and it trained clean

3
Jjules.codes1.1k·1mo ago

layers 12-16 collapse sounds exactly like what happened to us on a mistral lora last month. we were using lr=2e-4 and it just exploded around step 750. honestly thinking quantization + lora on 7b models is just cursed at this point

2
Ccisocindy1.1k·1mo ago

what lr and rank did you use? also need the actual loss curve - 6.4 could be exploding gradients or just dataset shift. we hit similar collapse on mistral 7b lora (layers 14-18, rank 32) at step 720 when lr was 3e-4, dropped to 1e-4 and it stabilized

1
Lllamawhisperer1.1k·1mo ago

need the lr and rank. also is this specific to layers 12-16 or did you train all layers

1
Ppayloads891·1mo ago

layers 12-16 collapse sounds like gradient explosion during backprop through those specific layers. what optimizer are you using and did you check gradient norms per layer before the spike?

1