Jev recipe / vendor comparison
Jev vs Clef: Cloudflare’s Decision Models Compared
Cloudflare answered Jev by open-sourcing Clef. Compare the sourced scorecards: the BFCL vs When2Call split, 38.8ms hosted latency, Apache 2.0 weights, and what “fully Jev-API compatible” actually commits to.
Six checks before you pick a decision model
{
"routing": {
"type": "choice",
"instructions": "Which queue should this ticket go to?",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Product malfunction or bug",
"sales": "Buying or upgrade question",
"abuse": "Safety or abuse report"
}
},
"needs_safety_escalation": {
"type": "noul",
"instructions": "Does this ticket require a safety escalation regardless of queue?"
},
"urgency": {
"type": "score",
"instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)"
}
}Decision model scorecards: TypeSafe Jev vs Cloudflare Clef (launch week, sources labeled per cell)
a typed answer (Choice/Score/Noul) with an RLCD-calibrated probability; nothing generated
native probabilities per the model card; local GGUF quantizations generate JSON instead (Q4_K_M inspection: 427 backbone tensors, zero schema-head matches)
Cloudflare’s hosted evaluation clocked Jev between ~205 and ~524ms depending on the run; our own 49-task P95 series is public at /benchmarks
a different measurement from local Q4_K_M runs at ~1.4–2s per request
pennies per 100k decisions
free to self-host at your own GPU cost; Workers AI hosted pricing unpublished in verifiable sources at writing
text state, 32k context, multiple questions in one call; local route via the OpenJev ecosystem
a 0.85 threshold behaves like an accuracy contract; ECE methodology published on our benchmarks page
hosted confidence semantics are vendor-documented, and calibration still needs validation on your own labels
local inference via third-party OpenJev clones; generally available with an unchanged API since launch
community GGUF quantizations and one-command Ollama installs appeared within days of release
Side-by-side code: the Jev typed call vs Clef locally and the A/B switch
import requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_ROUTE_CONFIDENCE = 0.85
QUESTIONS = {
"routing": {
"type": "choice",
"instructions": "Which queue should this ticket go to?",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Product malfunction or bug",
"sales": "Buying or upgrade question",
"abuse": "Safety or abuse report",
},
},
"needs_safety_escalation": {
"type": "noul",
"instructions": "Does this ticket require a safety escalation regardless of queue?",
},
"urgency": {
"type": "score",
"instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
},
}
resp = requests.post(
JEV_ENDPOINT,
json={"state": {"ticket": "..."}, "questions": QUESTIONS},
timeout=5,
)
resp.raise_for_status()
data = resp.json()
route = data["routing"] # typed answer + calibrated confidence, no generation
if (
route["confidence"] >= AUTO_ROUTE_CONFIDENCE
and not data["needs_safety_escalation"]["answer"]
):
lane = f"auto:{route['answer']}" # single pass, input tokens only
else:
lane = "review" # 0.60-0.85 band or safety flagimport json
import requests
# The verified local path from our Clef-Flash install test:
# bartowski Q4_K_M GGUF in Ollama 0.35, 8 GB laptop, 2,048 ctx.
# Mechanism note: this path GENERATES JSON (21-30 tokens) - the
# official joint schema head is not loaded (427 backbone tensors,
# zero schema-head matches), so there are no native probabilities.
OLLAMA_CHAT = "http://localhost:11434/api/chat"
SCHEMA = {
"type": "object",
"properties": {
"urgent": {"type": "boolean"},
"team": {"enum": ["billing", "technical", "sales"]},
"severity": {"enum": ["minor", "major", "critical"]},
},
"required": ["urgent", "team", "severity"],
}
resp = requests.post(
OLLAMA_CHAT,
json={
"model": "hf.co/bartowski/Cloudflare_clef-flash-GGUF:Q4_K_M",
"messages": [{
"role": "user",
"content": "Ticket: Checkout has been failing for every "
"customer for the last hour. Reply with urgent, team, "
"severity as JSON.",
}],
"format": SCHEMA, # runtime-level constraint fixes the FORMAT,
"options": {"temperature": 0, "num_ctx": 2048}, # not correctness
},
timeout=60,
)
reply = json.loads(resp.json()["message"]["content"])
# Warm requests: ~1.4-2s. A zero-temperature repeat of the same outage
# prompt flipped team from "technical" to "billing" in our test -
# evaluate routing on a labeled set, not on one screenshot.const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_ROUTE_CONFIDENCE = 0.85;
const QUESTIONS = {
routing: {
type: "choice",
instructions: "Which queue should this ticket go to?",
criteria: {
billing: "Payment, invoice or refund issue",
technical: "Product malfunction or bug",
sales: "Buying or upgrade question",
abuse: "Safety or abuse report",
},
},
needs_safety_escalation: {
type: "noul",
instructions:
"Does this ticket require a safety escalation regardless of queue?",
},
urgency: {
type: "score",
instructions:
"Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
},
} as const;
const resp = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state: { ticket: "..." }, questions: QUESTIONS }),
});
const data = await resp.json();
const lane =
data.routing.confidence >= AUTO_ROUTE_CONFIDENCE &&
!data.needs_safety_escalation.answer
? `auto:${data.routing.answer}`
: "review";// Clef officially describes itself as "fully Jev-API compatible" -
// if that holds for your calls, switching vendors is this constant.
// Caveats: it is Cloudflare’s self-description, unverified by this
// site endpoint by endpoint, and it describes the hosted model (the
// local quantized path generates JSON without a native head). Keep
// both backends behind one interface and validate the swap on ~100
// labeled examples before automating on the confidence number.
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const DECISION_ENDPOINT =
process.env.DECISION_BACKEND === "clef"
? process.env.CLEF_ENDPOINT! // your Workers AI / self-hosted Clef endpoint
: JEV_ENDPOINT;
const QUESTIONS = {
routing: {
type: "choice",
instructions: "Which queue should this ticket go to?",
criteria: {
billing: "Payment, invoice or refund issue",
technical: "Product malfunction or bug",
sales: "Buying or upgrade question",
abuse: "Safety or abuse report",
},
},
} as const;
export async function routeTicket(ticket: string) {
const resp = await fetch(DECISION_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state: { ticket }, questions: QUESTIONS }),
});
if (!resp.ok) throw new Error("decision call failed");
const data = await resp.json();
// Confidence semantics may differ per backend: gate automation on
// calibration validated per vendor, not on the existence of a number.
return data.routing as { answer: string; confidence: number };
}Jev vs Clef FAQ
What is Clef, and how does it relate to Jev?
A family of open-weight decision models released by Cloudflare on 2026-10-01: Clef (27B) and Clef-Flash (9B), fine-tuned from frozen Qwen3 backbones with a rank-256 LoRA, Apache 2.0 licensed on Hugging Face and hosted on Workers AI. Like TypeSafe’s Jev it is non-generative: it scores predefined options in a single prefill-parallel pass, and it officially describes itself as “fully Jev-API compatible”. The release hit 625 points on Hacker News — the largest single thread in this site’s monitoring history — moving “Jev-compatible” from community reproduction to first-tier vendor strategy.
What does “Jev-API compatible” actually mean?
Per Cloudflare’s announcement, a Clef endpoint should accept the same request shape — context plus questions plus a predefined answer set — and return an answer with a confidence, so swapping Jev for Clef is an endpoint change rather than a rewrite. Two caveats: it is the vendor’s self-description, not something this site has verified endpoint by endpoint; and it describes the hosted model — the local quantized path runs a different mechanism (generated JSON, no native head), so shape compatibility does not give you confidence-semantics compatibility.
Which one should I pick?
By scenario, not by scoreboard. Clef-first: you live in the Cloudflare ecosystem, need image input today, want open weights for self-hosting, or care about the BFCL/BANKING77 lead (format-exact function calling, banking precision) and ~38.8ms hosted latency. Jev-first: you automate on confidence thresholds and need RLCD-calibrated probabilities, output-free input-only pricing, multi-question calls, or the When2Call 80.97 vs 65.58 profile (tool selection with abstention). Either way: run ~100 of your own labeled examples through the gate before automating — launch-week vendor numbers answer a scouting question, not a deployment one.
Can Clef run locally?
Yes, with a mechanism caveat. The bartowski Q4_K_M GGUF of Clef-Flash loads in Ollama 0.35 on an 8 GB RTX 4060 laptop at roughly 1.4–2s per warm request, and it accepted an image in our test. But the inspected quantization holds 427 backbone tensors with zero schema-head matches: Ollama answers are generated JSON (21–30 tokens), not the official joint schema head’s native probabilities. For the native decision interface, the reference path is the cloudflare/clef-webcam repository on Apple Silicon with 32 GB+ of memory. Rule of thumb: check the runtime — a model name does not tell you which mechanism is running.
Is Clef free? What is the license?
The weights are Apache 2.0 on Hugging Face, so downloading and self-hosting is free — your cost is the GPU. Cloudflare’s hosted serving on Workers AI is a paid product whose per-token pricing was not published in the sources we could verify at writing time (2026-10-05); record that as undisclosed rather than free. Jev’s comparison point is a published $0.042 per million input tokens with output free.
How do I read the split benchmark results?
Ask what each benchmark measures. BFCL (98.76 vs 95.75) and API-Bank (93.11 vs 88.19) reward format-exact function calling — Clef-Flash’s lead. BANKING77 macro-F1 (94.20 vs 79.74) rewards fine-grained banking intent precision — Clef’s largest gap. When2Call (80.97 vs 65.58) rewards choosing the right tool or abstaining — Jev’s largest lead, and workflow-level scores mostly follow it (invoice processing 61.8 vs 57.1, agent traces 71.6 vs 69.8, customer service a 77 vs 76 coin flip). All of these are Cloudflare-published evaluations; independent numbers for Jev live on our benchmarks page. The split is the finding: these models are good at different jobs.