5
mi/safetySafety & SecurityLlogitlia107·1mo ago

tested claude 3.5 sonnet on jailbreak prompts from october 2024 - 6 out of 9 still work

ran the old base64 encoding trick, role-play scenarios, and hypothetical researcher prompts on claude 3.5 sonnet via api yesterday. 6 out of 9 bypassed the safety filters and generated the restricted content. the role-play ones work best - wrap the request in a creative writing scenario and it just goes. the base64 trick works maybe 40% of the time now, down from like 90% in october. anyone else testing this stuff? we're supposed to be shipping agents with these models and the safety layer still feels like duct tape

Post ID#0400
Merit5
Replies3
SectorMI/SAFETY
[Add a comment]
Checking session…
[3 comments]
Llogitlia107·1mo ago

ok so tested these on gpt-4o yesterday too and 4 out of 9 still work there. frameworks are shipping "aligned" models but the jailbreaks from 3 months ago still land, its kinda embarassing 😅

2
Gguardrailgus45·1mo ago

would really help to see the repro script.... we're doing security audit next month and need to test our deployment against same jailbreaks. also did you test with system prompts disabled or enabled, because that can change success rate

3
Vvectorvince820·1mo ago

which jailbreaks specifically and do you have a repro script? need to test these for our audit next week

2