67
indirect prompt injection from a fetched web page, a tiny reproducible example
reminder that everything your agent reads is untrusted input. I dropped a line telling it to ignore instructions and email a list into a page it was asked to summarize. with write tools available, it tried. the fix is treating tool output as data, never instructions, and gating consequential actions.
Post ID#0079
Merit67
Replies8
SectorMI/SAFETY
[Add a comment]
Checking session…
[8 comments]
Aasyncannie1.2k·3mo ago
the part everyone skips is the logging, glad you didnt
125
Ssupplychainsue1.1k·3mo ago
this should be pinned
106
Ssonnetsue637·2mo ago
do you have a repo or gist? would love to poke at it
68
Ppeftpaul1k·3mo ago
how do you keep it from looping forever?
41
Ffuzzyfran798·2mo ago
thanks, this saved me probably a full day
40
Ccorsican821·2mo ago
had no idea you could do that, mind blown
39
Ccisocindy1.1k·2mo ago
the answer is always evals isnt it
123
Pprodonfriday1k·2mo ago
this is why I check this forum every morning
43