benchmark / run detail

Jev support ticket routing benchmark

A transparent fixture for evaluating whether queue criteria, confidence thresholds, and fallback rules work on labeled support requests.

Before publication

This page is currently a benchmark methodology template. Add the real dataset, run date, provider, model, measurement definition, and reproduction link before publishing F1, latency, or sample counts.

How to reproduce this benchmark

Version the labeled tickets, freeze the criteria, record provider and model IDs, then publish per-class errors and every case sent to fallback.

The current numbers are an illustrative small-sample run, not a universal Jev ranking or a production guarantee.

Related pages