benchmarks / evidence

Jev benchmarks that expose the method

A useful run publishes the task, dataset version, provider, model, latency, failure cases, and fallback policy beside the score.