benchmarks / evidence
Jev benchmarks that expose the method
A useful run publishes the task, dataset version, provider, model, latency, failure cases, and fallback policy beside the score.
Jev support ticket routing benchmark
A transparent fixture for evaluating whether queue criteria, confidence thresholds, and fallback rules work on labeled support requests.
Evidence fields pending: publish a dataset version, run date, provider, and reproducible result before treating this as a measured run.
Jev spam detection benchmark
A task-level fixture for measuring spam decisions, false positives, latency, and the share of messages that require review.
Evidence fields pending: publish a dataset version, run date, provider, and reproducible result before treating this as a measured run.
Jev prompt injection detection benchmark
A small safety fixture for deciding whether an instruction can enter an agent workflow or must move to review.
Evidence fields pending: publish a dataset version, run date, provider, and reproducible result before treating this as a measured run.