jev-rerank-bench GitHub avatar by anessbelbati

anessbelbati /

Evaluation & Observability

jev-rerank-bench

Compares Jev, dedicated rerankers and chat models on the same retrieved passages.

License: MITLanguage: TypeScriptLast push: 9/17/2026#evaluation-benchmarks#search-retrieval#classification-ranking

What Jev Does in this Architecture

Ranks candidate passages with Choice, Noul and rubric scores, then computes retrieval metrics.

Key Architecture Benefit

Publishes raw responses, scoring code and per-dataset results for inspection.

Review & Benchmark Note: The author reports equal-dataset nDCG@10 of 0.692 for Jev and 0.691 for Cohere on eight English datasets, without establishing a winner. Query weighting changes the comparison. README and integration source reviewed at a fixed commit; not independently run or benchmarked by this site.
Clone Repository
git clone https://github.com/anessbelbati/jev-rerank-bench.git

Inspect on GitHub

Check out the upstream repository README, issues, and commit log.

anessbelbati/jev-rerank-bench

Architecture Category

Evaluation & Observability

Evaluation harnesses, latency profiling, tracing, and telemetry for Jev decisions.

View all in this category →

Related Projects in Evaluation & Observability

jev-review repository icon by NiazMorshed2007

NiazMorshed2007

jev-review

126
Evaluation & ObservabilityMITTypeScript

A local MCP code-quality reviewer returning structured scores to coding Agents.

What Jev does here

Jev scores correctness, complexity, tests and security; code ranks areas to improve.

Compares scores across checkpoints while leaving code changes to the primary Agent.

#coding-agents#mcp-integrations#evaluation-benchmarks
104 evidence links
Inspect details
jev-benchmarks repository icon by AbdelStark

AbdelStark

jev-benchmarks

9
Evaluation & ObservabilityApache-2.0Python

A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.

What Jev does here

Runs the same labeled text tasks through both backends and records probabilities, latency and failures.

Helps examine task-specific accuracy and whether confidence scores support chosen thresholds.

#evaluation-benchmarks#classification-ranking
3 evidence links
Inspect details
Evaluation & ObservabilityMITTypeScript

A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.

What Jev does here

Experiments send action candidates or typed questions to Jev, then execute or record the answers.

Includes source, experiment notes and some offline replays for comparing decision designs.

#evaluation-benchmarks#games-simulation#browser-automation
4 evidence links
Inspect details