benchmark / run detail

Jev spam detection benchmark

A task-level fixture for measuring spam decisions, false positives, latency, and the share of messages that require review.

Before publication

This page is currently a benchmark methodology template. Add the real dataset, run date, provider, model, measurement definition, and reproduction link before publishing F1, latency, or sample counts.

Measure the expensive mistakes

Keep ordinary messages, obvious spam, and ambiguous promotions in separate slices. Report false positives and review rate, not only aggregate F1.

This directional run is useful as a template. Replace the sample with your own traffic before setting an automation threshold.

Related pages