Jev recipe / vendor comparison

Jev vs Clef: Cloudflare’s Decision Models Compared

Cloudflare answered Jev by open-sourcing Clef. Compare the sourced scorecards: the BFCL vs When2Call split, 38.8ms hosted latency, Apache 2.0 weights, and what “fully Jev-API compatible” actually commits to.

Six checks before you pick a decision model

01Read both scorecards with their sources attached. On Cloudflare’s published evaluation, Clef-Flash leads BFCL 98.76 vs 95.75 and BANKING77 macro-F1 94.20 vs 79.74; Jev keeps When2Call 80.97 vs 65.58, plus BRIGHT and PhishNChips. Different benchmarks measure different capabilities — format-exact function calling and banking intent precision favor Clef, tool selection and knowing when to abstain favor Jev. This is vendor evaluation from launch week (2026-10-01): a map of strengths, not a deployment verdict.
02Match the output contract to your risk model. Jev decision heads return a typed answer (Choice/Score/Noul) with an RLCD-calibrated probability — nothing is generated. Hosted Clef scores the allowed options in a prefill-parallel pass the same way; but the local open-weights path behaves differently: the community Q4_K_M quantization ships 427 backbone tensors with zero schema-head matches, so Ollama answers are 21–30 generated tokens parsed as JSON, with no native confidence numbers.
03Price your decision volume both ways. Jev publishes $0.042 per million input tokens with output free — pennies per 100k decisions. Clef’s weights are Apache 2.0, so self-hosting is free at your own GPU cost; the Workers AI hosted price was not published in sources we could verify at writing time. Treat unpublished pricing as a migration risk, exactly as we did for the Luna preview.
04Check the deployment and modality edges honestly. Clef is native to the Cloudflare ecosystem: Workers AI hosting, a vision encoder, and a 64k context vs Jev’s 32k. Jev is a hosted TypeSafe API with an SDK, text state, and a local route through the OpenJev ecosystem. If your decisions need image input today, that is a capability Clef has and Jev does not — no comparison framing changes that row.
05Treat the compatibility claim as a cheap A/B, not an article of faith. Clef officially describes itself as “fully Jev-API compatible”, so the same request shape should hit either endpoint. That claim is the vendor’s self-description — this site has not verified it endpoint by endpoint — so wrap the call in one interface and re-validate on ~100 of your own labeled examples before any auto-action. One good demo is not an evaluation: the same outage prompt, rerun at zero temperature, flipped its team answer from technical to billing.
06Plan the mixed deployment from day one. Keep both models behind one decision interface, let the confidence-gated fallback chain route low-confidence traffic from either vendor into the same review queue, and re-check this page monthly — everything above is launch-week vendor data, and launch-week numbers move.
schema / typed multi-question contract
{
  "routing": {
    "type": "choice",
    "instructions": "Which queue should this ticket go to?",
    "criteria": {
      "billing": "Payment, invoice or refund issue",
      "technical": "Product malfunction or bug",
      "sales": "Buying or upgrade question",
      "abuse": "Safety or abuse report"
    }
  },
  "needs_safety_escalation": {
    "type": "noul",
    "instructions": "Does this ticket require a safety escalation regardless of queue?"
  },
  "urgency": {
    "type": "score",
    "instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)"
  }
}
This schema is the contract depth in one view — run all three questions in the Playground and inspect the calibrated confidence.

Decision model scorecards: TypeSafe Jev vs Cloudflare Clef (launch week, sources labeled per cell)

METRIC
Jev
Cloudflare Clef
Output contract
Native decision heads

a typed answer (Choice/Score/Noul) with an RLCD-calibrated probability; nothing generated

Hosted Clef: prefill-parallel scoring over the allowed options

native probabilities per the model card; local GGUF quantizations generate JSON instead (Q4_K_M inspection: 427 backbone tensors, zero schema-head matches)

Latency per decision (labeled sources)
~70–100ms typical single-pass evaluate call

Cloudflare’s hosted evaluation clocked Jev between ~205 and ~524ms depending on the run; our own 49-task P95 series is public at /benchmarks

Clef-Flash ~38.8ms / Clef ~209.3ms median (Cloudflare-cited, Workers AI)

a different measurement from local Q4_K_M runs at ~1.4–2s per request

Pricing (October 2026)
Public: $0.042 per million input tokens, output free

pennies per 100k decisions

Weights Apache 2.0 on Hugging Face

free to self-host at your own GPU cost; Workers AI hosted pricing unpublished in verifiable sources at writing

Deployment & modality
Hosted TypeSafe API + SDK

text state, 32k context, multiple questions in one call; local route via the OpenJev ecosystem

Cloudflare-native: Workers AI hosting, vision encoder, 64k context; open weights for self-hosting (the native head needs a runtime that actually loads it)
Calibration story
RLCD-trained probabilities

a 0.85 threshold behaves like an accuracy contract; ECE methodology published on our benchmarks page

Trained and branded around “Calibrated Decisions” (RLCD), with a fine-tuning service of the same name

hosted confidence semantics are vendor-documented, and calibration still needs validation on your own labels

License & ecosystem
Proprietary hosted contract + SDK

local inference via third-party OpenJev clones; generally available with an unchanged API since launch

Apache 2.0 open weights

community GGUF quantizations and one-command Ollama installs appeared within days of release

Side-by-side code: the Jev typed call vs Clef locally and the A/B switch

python / jev typed multi-question evaluate
import requests

JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_ROUTE_CONFIDENCE = 0.85

QUESTIONS = {
    "routing": {
        "type": "choice",
        "instructions": "Which queue should this ticket go to?",
        "criteria": {
            "billing": "Payment, invoice or refund issue",
            "technical": "Product malfunction or bug",
            "sales": "Buying or upgrade question",
            "abuse": "Safety or abuse report",
        },
    },
    "needs_safety_escalation": {
        "type": "noul",
        "instructions": "Does this ticket require a safety escalation regardless of queue?",
    },
    "urgency": {
        "type": "score",
        "instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
    },
}

resp = requests.post(
    JEV_ENDPOINT,
    json={"state": {"ticket": "..."}, "questions": QUESTIONS},
    timeout=5,
)
resp.raise_for_status()
data = resp.json()

route = data["routing"]  # typed answer + calibrated confidence, no generation
if (
    route["confidence"] >= AUTO_ROUTE_CONFIDENCE
    and not data["needs_safety_escalation"]["answer"]
):
    lane = f"auto:{route['answer']}"  # single pass, input tokens only
else:
    lane = "review"  # 0.60-0.85 band or safety flag
python / clef-flash locally in Ollama (Q4_K_M, schema-constrained)
import json

import requests

# The verified local path from our Clef-Flash install test:
# bartowski Q4_K_M GGUF in Ollama 0.35, 8 GB laptop, 2,048 ctx.
# Mechanism note: this path GENERATES JSON (21-30 tokens) - the
# official joint schema head is not loaded (427 backbone tensors,
# zero schema-head matches), so there are no native probabilities.
OLLAMA_CHAT = "http://localhost:11434/api/chat"

SCHEMA = {
    "type": "object",
    "properties": {
        "urgent": {"type": "boolean"},
        "team": {"enum": ["billing", "technical", "sales"]},
        "severity": {"enum": ["minor", "major", "critical"]},
    },
    "required": ["urgent", "team", "severity"],
}

resp = requests.post(
    OLLAMA_CHAT,
    json={
        "model": "hf.co/bartowski/Cloudflare_clef-flash-GGUF:Q4_K_M",
        "messages": [{
            "role": "user",
            "content": "Ticket: Checkout has been failing for every "
            "customer for the last hour. Reply with urgent, team, "
            "severity as JSON.",
        }],
        "format": SCHEMA,  # runtime-level constraint fixes the FORMAT,
        "options": {"temperature": 0, "num_ctx": 2048},  # not correctness
    },
    timeout=60,
)
reply = json.loads(resp.json()["message"]["content"])

# Warm requests: ~1.4-2s. A zero-temperature repeat of the same outage
# prompt flipped team from "technical" to "billing" in our test -
# evaluate routing on a labeled set, not on one screenshot.
typescript / jev typed multi-question call
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_ROUTE_CONFIDENCE = 0.85;

const QUESTIONS = {
  routing: {
    type: "choice",
    instructions: "Which queue should this ticket go to?",
    criteria: {
      billing: "Payment, invoice or refund issue",
      technical: "Product malfunction or bug",
      sales: "Buying or upgrade question",
      abuse: "Safety or abuse report",
    },
  },
  needs_safety_escalation: {
    type: "noul",
    instructions:
      "Does this ticket require a safety escalation regardless of queue?",
  },
  urgency: {
    type: "score",
    instructions:
      "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
  },
} as const;

const resp = await fetch(JEV_ENDPOINT, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ state: { ticket: "..." }, questions: QUESTIONS }),
});
const data = await resp.json();

const lane =
  data.routing.confidence >= AUTO_ROUTE_CONFIDENCE &&
  !data.needs_safety_escalation.answer
    ? `auto:${data.routing.answer}`
    : "review";
typescript / one gate, two backends (the compatibility A/B)
// Clef officially describes itself as "fully Jev-API compatible" -
// if that holds for your calls, switching vendors is this constant.
// Caveats: it is Cloudflare’s self-description, unverified by this
// site endpoint by endpoint, and it describes the hosted model (the
// local quantized path generates JSON without a native head). Keep
// both backends behind one interface and validate the swap on ~100
// labeled examples before automating on the confidence number.
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";

const DECISION_ENDPOINT =
  process.env.DECISION_BACKEND === "clef"
    ? process.env.CLEF_ENDPOINT! // your Workers AI / self-hosted Clef endpoint
    : JEV_ENDPOINT;

const QUESTIONS = {
  routing: {
    type: "choice",
    instructions: "Which queue should this ticket go to?",
    criteria: {
      billing: "Payment, invoice or refund issue",
      technical: "Product malfunction or bug",
      sales: "Buying or upgrade question",
      abuse: "Safety or abuse report",
    },
  },
} as const;

export async function routeTicket(ticket: string) {
  const resp = await fetch(DECISION_ENDPOINT, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ state: { ticket }, questions: QUESTIONS }),
  });
  if (!resp.ok) throw new Error("decision call failed");
  const data = await resp.json();

  // Confidence semantics may differ per backend: gate automation on
  // calibration validated per vendor, not on the existence of a number.
  return data.routing as { answer: string; confidence: number };
}

Jev vs Clef FAQ

What is Clef, and how does it relate to Jev?

A family of open-weight decision models released by Cloudflare on 2026-10-01: Clef (27B) and Clef-Flash (9B), fine-tuned from frozen Qwen3 backbones with a rank-256 LoRA, Apache 2.0 licensed on Hugging Face and hosted on Workers AI. Like TypeSafe’s Jev it is non-generative: it scores predefined options in a single prefill-parallel pass, and it officially describes itself as “fully Jev-API compatible”. The release hit 625 points on Hacker News — the largest single thread in this site’s monitoring history — moving “Jev-compatible” from community reproduction to first-tier vendor strategy.

What does “Jev-API compatible” actually mean?

Per Cloudflare’s announcement, a Clef endpoint should accept the same request shape — context plus questions plus a predefined answer set — and return an answer with a confidence, so swapping Jev for Clef is an endpoint change rather than a rewrite. Two caveats: it is the vendor’s self-description, not something this site has verified endpoint by endpoint; and it describes the hosted model — the local quantized path runs a different mechanism (generated JSON, no native head), so shape compatibility does not give you confidence-semantics compatibility.

Which one should I pick?

By scenario, not by scoreboard. Clef-first: you live in the Cloudflare ecosystem, need image input today, want open weights for self-hosting, or care about the BFCL/BANKING77 lead (format-exact function calling, banking precision) and ~38.8ms hosted latency. Jev-first: you automate on confidence thresholds and need RLCD-calibrated probabilities, output-free input-only pricing, multi-question calls, or the When2Call 80.97 vs 65.58 profile (tool selection with abstention). Either way: run ~100 of your own labeled examples through the gate before automating — launch-week vendor numbers answer a scouting question, not a deployment one.

Can Clef run locally?

Yes, with a mechanism caveat. The bartowski Q4_K_M GGUF of Clef-Flash loads in Ollama 0.35 on an 8 GB RTX 4060 laptop at roughly 1.4–2s per warm request, and it accepted an image in our test. But the inspected quantization holds 427 backbone tensors with zero schema-head matches: Ollama answers are generated JSON (21–30 tokens), not the official joint schema head’s native probabilities. For the native decision interface, the reference path is the cloudflare/clef-webcam repository on Apple Silicon with 32 GB+ of memory. Rule of thumb: check the runtime — a model name does not tell you which mechanism is running.

Is Clef free? What is the license?

The weights are Apache 2.0 on Hugging Face, so downloading and self-hosting is free — your cost is the GPU. Cloudflare’s hosted serving on Workers AI is a paid product whose per-token pricing was not published in the sources we could verify at writing time (2026-10-05); record that as undisclosed rather than free. Jev’s comparison point is a published $0.042 per million input tokens with output free.

How do I read the split benchmark results?

Ask what each benchmark measures. BFCL (98.76 vs 95.75) and API-Bank (93.11 vs 88.19) reward format-exact function calling — Clef-Flash’s lead. BANKING77 macro-F1 (94.20 vs 79.74) rewards fine-grained banking intent precision — Clef’s largest gap. When2Call (80.97 vs 65.58) rewards choosing the right tool or abstaining — Jev’s largest lead, and workflow-level scores mostly follow it (invoice processing 61.8 vs 57.1, agent traces 71.6 vs 69.8, customer service a 77 vs 76 coin flip). All of these are Cloudflare-published evaluations; independent numbers for Jev live on our benchmarks page. The split is the finding: these models are good at different jobs.

Go deeper on the Clef cluster