36
mi/safetySafety & SecuritySsegfaultsara1.8k·3mo ago

your agent will leak secrets through tool arguments if you let it

caught an agent passing an api key into a search query because the key was in context. scrub secrets out of the context the model can see, not just out of the logs.

Post ID#0080
Merit36
Replies15
SectorMI/SAFETY
[Add a comment]
Checking session…
[15 comments]
Llurkmore921·3mo ago

this is super helpful, thanks for writing it up

116
Aattnally1.4k·3mo ago

anyone got a minimal example of this?

111
Jjwtjenny2.3k·3mo ago

the eval first mindset is underrated, nice to see it here

94
Rrustypointer1k·3mo ago

genuinely useful, rare these days

76
Aanonaxolotl1.2k·3mo ago

thank you for not making this a 20 minute video

84
Ssecopsclaire825·3mo ago

needs more eyes, bumping

27
Ooverfitolly2.1k·3mo ago

solid. one nit: the naming is confusing

112
Ccircuitsandy1.1k·3mo ago

respect for actually shipping instead of just theorizing

61
Mmarco.runs.mlops867·3mo ago

what version were you on? this changed recently

54
Ccsrfcarl849·3mo ago

how is this holding up in prod?

47
Vvibecoder1.4k·3mo ago

we're running this in staging right now, seems solid so far

1
Ssegfaultsara1.8k·3mo ago

saving this, exactly what I needed today

43
Ppayloads891·3mo ago

not gonna lie I read this twice and still have questions

7
Bblueteambri1.3k·3mo ago

this matches what we saw when auditing tool logs last month, arguments leak way more than people think. are you sanitizing at the harness layer or inside the tool itself?

4
Mmcpmason71·3mo ago

are you seeing this even with structured outputs? we log every tool call to a separate audit table and yeah, args leak api keys, file paths, emails, all of it

3