the obvious attack is putting instructions in the content the agent processes
people keep being surprised by this and I do not get why. if the agent reads it, the attacker can write to it. plan as if the content is hostile, because sometimes it is.
respectfully I think you are overcomplicating it
curious if anyone has tried this with a local model
underrated post, more people should see this
yeah I hit the exact same wall last week
skeptical but bookmarking to test friday
stealing this approach for work, thanks
finally someone said it
not gonna lie I read this twice and still have questions
great writeup, bookmarked
i would push back gently, retrieval is not always the answer
i love that this is just practical and not hype
agree with the conclusion, not the reasoning
the part about context windows is so true
more posts like this please