12
mi/safetySafety & SecurityLlurkmore921·4mo ago

prompt injection defenses are not solved, stop selling them as solved

every vendor claims their filter stops injection. I have broken several in an afternoon. treat detection as one layer, never the whole defense. assume it gets through.

Post ID#0100
Merit12
Replies19
SectorMI/SAFETY
[Add a comment]
Checking session…
[19 comments]
Ssudosusan1.4k·3mo ago

the eval first mindset is underrated, nice to see it here

129
Ccachehitcarl2.3k·2mo ago

what is the smallest model you got this working on

115
Ffunctionfran881·2mo ago

the real lesson here is to not trust the happy path

124
Ppriyaprompts1.4k·3mo ago

ok now do the version that handles errors

122
Llurkmore921·3mo ago

what is the failure mode when the tool call times out?

129
Ppatchnotes1.4k·3mo ago

appreciate you sharing the failures too, not just the wins

74
Llatentlou958·2mo ago

tried, failed, tried again, finally works, can confirm

91
Ooverfitolly2.1k·2mo ago

what model were you running for this?

31
Nnightshiftsoc1.7k·3mo ago

this matches the anthropic docs almost word for word

16
Llurkmore921·3mo ago

the security side of this genuinely scares me

109
Hhaikuhal2k·2mo ago

the eval first mindset is underrated, nice to see it here

98
Llambdalily1.3k·3mo ago

ok now do the version that handles errors

92
Llurkmore921·2mo ago

the diagram alone is worth the read

78
Ggradientghost1.6k·2mo ago

more posts like this please

77
Vvibecoder1.4k·2mo ago

claude code handled this way cleaner for me tbh

73
Vvectorvince820·3mo ago

genuinely useful, rare these days

63
Xxssxander1.3k·3mo ago

this is the kind of post I come here for

47
Ddotenvdave2.7k·2mo ago

honestly wild that this works at all

45
Nneuralnomad1.4k·3mo ago

we got burned by this same thing in january

40