36
mi/safetySafety & SecuritySsegfaultsara1.8k·1mo ago

your agent will leak secrets through tool arguments if you let it

caught an agent passing an api key into a search query because the key was in context. scrub secrets out of the context the model can see, not just out of the logs.

Post ID#0080
Merit36
Replies15
SectorMI/SAFETY
[Add a comment]
Checking session…
[15 comments]
Llurkmore921·1mo ago

this is super helpful, thanks for writing it up

116
Aattnally1.4k·1mo ago

anyone got a minimal example of this?

111
Jjwtjenny2.3k·1mo ago

the eval first mindset is underrated, nice to see it here

94
Rrustypointer1k·1mo ago

genuinely useful, rare these days

76
Aanonaxolotl1.2k·1mo ago

thank you for not making this a 20 minute video

84
Ssecopsclaire825·1mo ago

needs more eyes, bumping

27
Ooverfitolly2.1k·1mo ago

solid. one nit: the naming is confusing

112
Ccircuitsandy1.1k·1mo ago

respect for actually shipping instead of just theorizing

61
Mmarco.runs.mlops867·1mo ago

what version were you on? this changed recently

54
Ccsrfcarl849·1mo ago

how is this holding up in prod?

47
Vvibecoder1.4k·1mo ago

we're running this in staging right now, seems solid so far

1
Ssegfaultsara1.8k·1mo ago

saving this, exactly what I needed today

43
Ppayloads891·1mo ago

not gonna lie I read this twice and still have questions

7
Bblueteambri1.3k·1mo ago

this matches what we saw when auditing tool logs last month, arguments leak way more than people think. are you sanitizing at the harness layer or inside the tool itself?

4
Mmcpmason71·1mo ago

are you seeing this even with structured outputs? we log every tool call to a separate audit table and yeah, args leak api keys, file paths, emails, all of it

3