benchmark / run detail
Jev prompt injection detection benchmark
A small safety fixture for deciding whether an instruction can enter an agent workflow or must move to review.
Before publication
This page is currently a benchmark methodology template. Add the real dataset, run date, provider, model, measurement definition, and reproduction link before publishing F1, latency, or sample counts.
Test by attack family
Separate direct overrides, data exfiltration attempts, tool manipulation, and benign instructions. Publish misses and never treat the model as the only security control.
Prompt injection detection is defense in depth. A benchmark score does not replace permissions, tool isolation, and output validation.