Guides / illustrated walkthrough
Clef-Flash vs Jev: I Tested Cloudflare's Decision Model Locally in Ollama (Q4_K_M)
Cloudflare released Clef as open weights, so the practical question became testable: does the small quantized version actually run and decide correctly in Ollama on an 8 GB laptop? Two demos, one honest surprise about the joint schema head, and the published Clef-vs-Jev numbers kept carefully separate.
Quick takeaway
Cloudflare's Clef release made the obvious question testable: can the small version actually decide on your hardware? This video answers it with unusual care. The setup: bartowski's Clef-Flash GGUF in Q4_K_M, Ollama 0.35.0, an RTX 4060 laptop with 8 GB of VRAM, a 2,048-token context. Demo one routes a checkout-outage ticket: plain prompting returned malformed JSON (cold: 18.29 s including 15.90 s of loading), an Ollama JSON-schema constraint fixed the format, the first constrained run answered technical/urgent/critical at 2.27 s — and a zero-temperature repeat of the same prompt answered billing at 1.96 s. Demo two feeds a cropped still image and gets gesture=peace_sign, holding_drink=false at 1.35 s. Then the finding that reframes everything: the official release ships a separate joint schema head that scores the allowed options, but the local quantized file holds 427 backbone tensors with zero schema/head matches — the Ollama answers are 21–30 generated tokens parsed as JSON, not native decision probabilities. The published comparison (API-Bank 93.11 vs 88.19 for Clef-Flash; When2Call 80.97 vs 65.58 for Jev; mixed workflow scores) comes from Cloudflare's hosted evaluation and answers a different question than this laptop. Verdict: it runs locally; the native head stays separate; check the runtime, because the model name alone does not tell you which mechanism is running.
Video source
Prompt Engineer 48
Step-by-step walkthrough
- 1
The question: does the small Clef actually decide locally?
Cloudflare released Clef as a family of decision models with open weights, which turns a marketing question into a testable one: take the small version, put it into Ollama on an ordinary laptop, and see if it actually works. The frame lists what Clef-Flash adds per the official model card: a 9B Qwen3 backbone, vision inputs with a projector, and — the detail this whole page turns on — a separate joint schema scoring head. The plan is deliberately narrow: one text demo (support routing), one image demo, and a comparison against the original Jev using published evidence only. No benchmark theater; the video spends its time on the few examples it actually ran.

Per the model card: 9B Qwen3 backbone, vision inputs, and a separate joint schema head — remember that head.Watch at 3:00 - 2
The test machine: 8 GB of VRAM and honest constraints
The rig is deliberately ordinary: an NVIDIA RTX 4060 laptop GPU with 8 GB of VRAM, Ollama 0.35.0, the Q4_K_M quantization, and a 2,048-token context with a short output limit. The creator is upfront that this is a laptop run with modest context — it does not test a long context window and says nothing about production workloads. During the test, Ollama reported a mix of CPU and GPU execution with about 87% of compute on the GPU. Those constraints are the point: if the quantized package cannot decide on this class of hardware, the open-weights story has a hole in it.

RTX 4060 laptop, 8 GB VRAM, Ollama 0.35.0, Q4_K_M, 2,048-token context — deliberately ordinary hardware.Watch at 3:25 - 3
The local route: one ollama run command
Getting the model is the easy part and the frame shows all of it: ollama run hf.co/bartowski/Cloudflare_clef-flash-GGUF:Q4_K_M, after which requests go to the local Ollama endpoint at localhost:11434 with the 2,048 context setting. Two size numbers travel with this model and the video keeps them distinct: the quantization card lists the backbone file at about 5.84 GB, while Ollama reports roughly 6.8 GB for the installed package — the difference being the visual projector that ships alongside. The creator also notes honestly that the model was already present when recording started; what you see are inference checks, not a download.

One command to local: hf.co/bartowski/Cloudflare_clef-flash-GGUF:Q4_K_M on localhost:11434.Watch at 4:05 - 4
Demo 1: routing a checkout outage
The support ticket is deliberately short so nothing hides behind a big prompt: "Checkout has been failing for every customer for the last hour." Three fields are requested — urgent as yes/no, team from billing/technical/sales, severity on minor/major/critical — and the frame states the intended answers: urgent true, technical (a widespread service failure), critical. Two questions stay separate throughout the video: did the output follow the format, and did the selected answer make sense? Passing the first never settles the second.

One sentence in, three typed fields out — with the intended answers declared up front.Watch at 4:57 - 5
Plain prompting fails the format
The first attempt just asks for JSON in the prompt. The exact output: {"urgent": true, team technical, critical} — the values are semantically right but the string is not valid JSON, with keys and values missing required structure. Cold, the request took 18.29 seconds total, 15.90 of which was model loading. For any application expecting a parseable object, this is a real failure — and it is the same failure mode every LLM-with-JSON-prompt integration knows. Keep it in mind when the schema fix arrives two cards later.

Right idea, broken format: invalid JSON, 18.29 s cold (15.90 s of it loading).Watch at 5:25 - 6
The fix: a schema constraint at the runtime level
Instead of asking harder, the second attempt constrains the format in Ollama itself: urgent defined as a boolean, team as an enum over billing/technical/sales, severity as an enum over minor/major/critical. With the answer space enforced by the runtime, the output parses cleanly. The video credits this correctly — it demonstrates the benefit of constraining format at the runtime level rather than hoping the model hits it — and immediately flags that format was never the interesting question. Decision quality is still open.

Urgent: boolean. Team and severity: enums. The runtime now guarantees the shape.Watch at 5:45 - 7
Run A: the plausible answer — technical, 2.27 s
The first schema-constrained request returns exactly what the demo intended: team "technical", urgent true, severity "critical" — valid JSON at 2.27 seconds of Ollama wall time on the warm request. As the observation card says: plausible routing. If this were the only output shown, the experiment would look like a clean success and the video would be a victory lap. The creator keeps it on screen precisely so the next frame can take it away.

Run A: technical / true / critical, valid JSON, 2.27 s — exactly the intended answer.Watch at 6:05 - 8
Run B: the repeat disagreed — billing, 1.96 s
Same outage prompt, same allowed fields, zero temperature — and the repeat returns team "billing" at 1.96 seconds, keeping urgent true and severity critical. Both results stay in the video. The honest reading, which the narration spells out: one repeat cannot diagnose a cause, but this tested path did not give stable team selection, and billing is a perfectly valid department even when the message should have gone to technical. For an application that auto-routes every ticket, that instability is exactly the kind of thing a one-good-screenshot evaluation hides — you evaluate on a meaningful set of examples, not a demo.

Run B: same prompt, zero temperature — billing. The repeat is the whole review.Watch at 6:42 - 9
Demo 2: vision — peace sign at 1.35 s
The second demo exercises the visual input the hosted Clef-Flash advertises. The creator crops a still image so the model sees a hand gesture rather than a whole desktop, fixes the allowed gestures before the request, and asks the two questions: which gesture, and is the person holding a drink? The constrained answer: gesture "peace_sign", holding_drink false — at 1.35 seconds — and both check out against the input: two extended fingers, no drink. Useful evidence that the quantized package accepts images, with the video's own caveat attached: it is one picture, routed through generated JSON, not a vision benchmark.

One cropped still: peace_sign, holding_drink false, 1.35 s — the vision path works through Ollama.Watch at 8:05 - 10
The inference gap: the decision head never loaded
Now the finding that reframes both demos. The official release lists separate files — backbone.safetensors plus joint_head.safetensors and joint_schema_model.py — and the reference loader explicitly loads that head alongside the backbone, using it to score the allowed options. The inspected Ollama path tells a different story: 427 backbone tensors, zero schema/head name matches, and — per the request logs — generated output tokens. The support answers used around 30 tokens, the image answer 21. Those are generated strings parsed as JSON, not native option probabilities, and the video refuses to dress them up: "I will not invent confidence bars to make the demo look like the original decision interface."

Official release: backbone + joint head, loaded and scored. Ollama path: 427 backbone tensors, 0 head matches — generated JSON instead.Watch at 9:05 - 11
The published comparison: Clef-Flash vs Jev, labeled
Comparing against Jev switches evidence sources — these bars come from Cloudflare's published evaluation, not the laptop. Clef-Flash 93.11 vs original Jev 88.19 on API-Bank; When2Call flips it hard, 80.97 for Jev vs 65.58. The card carries its own warning label — vendor evaluation, not our Q4_K_M scores — and the narration adds the honest frame: published numbers decide what to investigate, not what to deploy. What they do establish is that the category now has two serious implementations with different strengths.

API-Bank to Clef-Flash (93.11 vs 88.19); When2Call to Jev (80.97 vs 65.58) — vendor evaluation, clearly labeled.Watch at 10:22 - 12
The workflow results are mixed — on purpose
The workflow-level published scores refuse a clean winner too: customer service exact actions 77 for Clef-Flash vs 76 for Jev (a coin flip), invoice processing 57.1 vs 61.8 to Jev, agent traces 69.8 vs 71.6 to Jev. The video draws the right boundary twice: these are reported evaluation scores that certify nothing about this quantization, and the latency numbers live in different worlds — Cloudflare reports ~38.8 ms median for hosted Clef-Flash against ~204.7 ms for Jev in that same hosted evaluation, while this laptop's Q4_K_M requests took about 1.4–2 s. Hosted vendor latency and local generated-JSON latency are different measurements; putting them in one chart would erase the difference.

77 vs 76, 57.1 vs 61.8, 69.8 vs 71.6 — no clean winner, and every number labeled by where it was produced.Watch at 10:45 - 13
The verdict: it runs locally. The native head stays separate.
The closing cards hold both truths at once. It runs locally: the Q4_K_M package works on an 8 GB laptop, accepted an image, and produced schema-valid JSON — with support routing that was mixed across repeated outage requests. Native head separate: Ollama output here is generated JSON, not the official decision mechanism. The takeaway generalizes past this model: check the runtime, because a model name does not identify the inference mechanism. And the practical path forward matches our own playbook — keep the explicit schema, validate the selected fields, review ambiguous or consequential cases (the outage repeat is the standing reason), and if you want the full decision-model experience, the next step is a runtime that actually loads the joint schema head — asking a generated answer to include a confidence number would not fix that gap.

It runs locally. Native head separate. — the two-sentence honest verdict this category needs more of.Watch at 12:25
Frequently asked questions
Can Cloudflare's Clef-Flash run locally in Ollama?
Yes — the bartowski Q4_K_M GGUF loads and responds on an 8 GB RTX 4060 laptop through Ollama 0.35.0 with a 2,048-token context, at roughly 1.4–2 s per warm request (87% GPU execution). The backbone file is about 5.84 GB and the installed package about 6.8 GB including the visual projector. What the video did not establish is that Ollama runs the native decision mechanism — see the joint schema head answer below.
What is the joint schema head — and why does it matter?
It is the part of the official Clef release that scores the allowed options to produce native typed probabilities: separate files (joint_head.safetensors, joint_schema_model.py) that the reference loader explicitly loads alongside the backbone. The video inspected the local quantized file and found 427 backbone tensors with zero schema/head name matches, and Ollama's logs showed generated output tokens (21–30 per answer) parsed as JSON. In other words, the local path generates decision-shaped text instead of running the decision head — which is why its answers carry no native confidence numbers.
Is Clef-Flash better than Jev?
The published Cloudflare evaluation splits: Clef-Flash leads API-Bank 93.11 to 88.19 while Jev leads When2Call 80.97 to 65.58; workflow scores are customer service 77 vs 76, invoice processing 57.1 vs 61.8 to Jev, and agent traces 69.8 vs 71.6 to Jev. Those are hosted vendor-evaluation numbers answering a different question than a local Q4_K_M run — treat them as a map of strengths, not a verdict, and test the workflow you actually ship.
Why did the same outage prompt return different teams?
The video ran the checkout-outage ticket twice with a JSON-schema constraint and zero temperature: run A answered technical at 2.27 s, run B answered billing at 1.96 s (both kept urgent true and severity critical). One repeat cannot establish a cause, but it shows the tested path did not give stable team selection — a valid option can still be the wrong decision. That is the argument for evaluating routing on a meaningful labeled set before letting any model auto-route tickets, and for reviewing ambiguous or consequential cases.
How do you get valid JSON from Clef-Flash in Ollama?
Constrain the format at the runtime level instead of asking in the prompt. Plain prompting returned malformed JSON ({'urgent': true, team technical, critical}) on the video's cold run; supplying an Ollama JSON schema — urgent as a boolean, team and severity as enums — made every subsequent output parse. Keep the explicit schema, validate the selected fields in code, and treat the constraint as a format guarantee, not a correctness guarantee.
What is the difference between this and the clef-webcam repo?
Execution path. The cloudflare/clef-webcam repository uses the original release with its reference inference code to produce typed decisions with probabilities from webcam frames — its quick start targets Apple Silicon with 32 GB+ of memory. The video's Windows experiment runs the community GGUF quantization through Ollama, which generates JSON instead of invoking the decision head. Both are legitimate, but they are different mechanisms wearing the same model name — the webcam repo is the one that demonstrates the native decision interface.
Related guides
Jev vs Luna: the Third-Party Benchmark Guide
The cited-numbers sibling of this page: a 505-sample independent evaluation of Jev and Luna. This page installs and tests; that one reads the scoreboard.
ReadJev vs OpenAI Decisions API
The product-API comparison axis: how the hosted decision contracts differ on output shape, latency, and confidence semantics.
ReadOllama Decision Models: tev1 and Nimble
The local decision-model setup that actually serves native typed outputs on consumer hardware — the control group for this experiment.
ReadOur Own 49-Task Benchmark
Independent F1, P95 latency, and calibration (ECE) measurements for Jev on real workloads — the same honesty standard this video holds Clef to.
ReadMore video walkthroughs
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Jev Log Triage with Expanso Edge
- Jev Lead Enrichment with Treg: ICP and Signup Scoring
- Use Jev Decision Nodes in Heym for Model Routing
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)
- Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)
- Jev Context Compaction: Prune AI Agent Memory Without Generative Summaries
- Jev as an LLM Judge: Confidence-Gated Cascades at 0.36% of the Cost
- TypeSafe Computer Use: Local Desktop Automation with Jev, Step by Step
- Jev Resume Screening: Build an AI Resume Evaluator with the Jev JavaScript SDK
- Jev + Claude Code Guide: Voice-Controlled Browser Automation with Typed Decisions
- Jev + Codex: Install the TypeSafe Skill and Triage a Real Gmail Inbox
- Jev vs Ollama: Can Local AI Replace Hosted Jev Without Sending Your Data Away?
- Build Your Own Jev: Train a Free Open-Source Zero-Shot Classifier (That Plays Doom)
- CUA-S1-Forms: a 706K-Parameter Jev-Like Model That Fills GUI Forms on Your CPU
- NOC/SOC Alert Triage with Jev: Rules First, One Typed Question, a Policy Gate
- Ollama Decision Models: Run tev1 and Nimble Locally (Tested on an 8 GB Card)
- Jev Guardrails in Production: A Five-Step Playbook for Decision Automation
- A Session Drift Guard for Pi Agent: Let the Jev Model Propose, Let Code Decide
- 50 Tev1 Use Cases: What a Local Decision Model Can Actually Do (Tested on an 8 GB Laptop)
- Jev, Hands-On: Where the Official Claims Meet Independent Remeasurement
- Jev Ticket Classification in a Real App: the After-Insert Hook and the Calculated Field
- OpenJev RLCD: Run the Open-Source Calibrated Decision Model Locally (Full Guide)