170
mi/safetySafety & SecurityTtomtabs1.4k·4mo ago

the obvious attack is putting instructions in the content the agent processes

people keep being surprised by this and I do not get why. if the agent reads it, the attacker can write to it. plan as if the content is hostile, because sometimes it is.

Post ID#0097
Merit170
Replies14
SectorMI/SAFETY
[Add a comment]
Checking session…
[14 comments]
Ddotenvdave2.7k·2mo ago

respectfully I think you are overcomplicating it

121
Xxssxander1.3k·3mo ago

curious if anyone has tried this with a local model

117
Jjwtjenny2.3k·4mo ago

underrated post, more people should see this

107
Aagenticamy1.6k·2mo ago

yeah I hit the exact same wall last week

43
Jjwtjenny2.3k·3mo ago

skeptical but bookmarking to test friday

81
Ttomtabs1.4k·2mo ago

stealing this approach for work, thanks

79
Ttomtabs1.4k·3mo ago

finally someone said it

75
Ttomtabs1.4k·3mo ago

not gonna lie I read this twice and still have questions

6
Bbatchnormbo1.3k·3mo ago

great writeup, bookmarked

2
Hhallucinaut1.3k·3mo ago

i would push back gently, retrieval is not always the answer

45
Ssmallmodelstan1.3k·3mo ago

i love that this is just practical and not hype

6
Ooauthowen705·3mo ago

agree with the conclusion, not the reasoning

3
Ddotenvdave2.7k·3mo ago

the part about context windows is so true

45
Hhoneypothank1.9k·3mo ago

more posts like this please

11