Guides / illustrated walkthrough

Jev vs Ollama: Can Local AI Replace Hosted Jev Without Sending Your Data Away?

A step-by-step breakdown of CloudAICode’s 14:26 comparison: hosted Jev at $0.042 per million input tokens versus local decision models — Kev on Qwen 3.5, SemIf score reading, Von on CPU — covering the accuracy gap, latency, calibration, and when keeping data on your own hardware is worth it.

Quick takeaway

CloudAICode’s 14:26 comparison asks one question: can typed, probability-backed AI decisions stay on your own hardware? The video opens with the cloud baseline — hosted Jev at $0.042 per million input tokens, about $0.0004 per case, with a 70–500 ms output chip — and then names the real blocker: when state holds customer emails, patient records, internal documents or financial details, “the cheapest cloud decision can still be impossible.” What you would replace is not a chat model: a typed decision must pick refund, escalate, reply or ignore and structurally cannot return a fifth option, which is why an Ollama-served chat LLM is not a drop-in Jev substitute — you can prompt it for JSON, but you lose the decision heads and the calibrated probabilities. The local route demonstrated is Kev, Jared Palmer’s Apache-2.0 adapters on Alibaba’s Qwen 3.5: serve it with KEV_DTYPE=bf16 uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009, then change one SDK line from the TypeSafe cloud address to http://127.0.0.1:8009 — “same System One interface, only the address changes.” In the video’s comparison run the new-source test lands at 68.4% for the 0.8B, 83.7% for the 4B and 85.2% for the 9B, against hosted Jev’s 85.7% (dev); calibration improves with size (Brier score 0.460 for the 0.8B versus 0.237 for the 9B on the test split); and on an Apple M5 one decision takes 329 ms (0.8B), 779 ms (4B) or roughly 2,000 ms (9B). SemIf (Theo Lee) skips training entirely by score-reading a compressed Qwen 3.5 4B you already host: 75.1% modal agreement against hosted Jev’s published 78.5% on the same 102-row, 20-case subset. Von runs on a CPU with no graphics card, and OpenJev adds visual judgment over images. The closing frames are a balance sheet: local gains are no metered token charge and context staying in your controlled environment; your problems become hardware and electricity, deployment and monitoring, calibration and security. Frequent, fixed answers that tolerate review — support routing, spam filtering, feedback tagging, lead scoring — are the local sweet spot, while the case weakens when the output cannot be listed beforehand, on a spectrum from support queues down to medical care. Because most published benchmarks come from the projects themselves on different datasets, the video’s final instruction is to test roughly 100 examples from your own workload and verify calibration before trusting any model with important decisions.

Video source

CloudAICode

14:26jxnywXprJQw

Step-by-step walkthrough

  1. 1

    Price the cloud baseline: $0.042 per million input tokens

    The comparison starts with the hosted card, not the local one. Jev on TypeSafe’s servers costs $0.042 per million input tokens — about $0.0004 per case — with the output chip showing a 70–500 ms decision. At four hundredths of a cent per call, the video’s point is that cost is almost never the reason to leave the cloud. Whatever pushes you local will be something else, and the next card names it.

    Jev hosted pricing card showing 0.042 dollars per million input tokens, about 0.0004 dollars per case, and an output chip reading 70 to 500 milliseconds
    The cloud baseline is nearly free — price alone will not send you local.Watch at 1:02
  2. 2

    Find the real blocker: privacy rules, not price

    The second card splits the screen. On the left, your organization holds customer emails, patient records, internal documents and financial details. On the right, hosted Jev sits on TypeSafe’s servers — nearly free, and marked “Unusable” the moment a red state-cannot-leave barrier stands between the two. This is the honest center of any Jev vs Ollama-style local debate: the cheapest cloud decision can still be impossible, because no accuracy or price number survives a compliance rule that forbids the data from leaving.

    Your organization card listing customer emails, patient records, internal documents and financial details separated from hosted Jev by a state-cannot-leave barrier and marked unusable
    When state cannot leave the building, even a nearly free decision is impossible.Watch at 1:22
  3. 3

    Know what you are replacing: a typed decision, not a chat

    Before reaching for a local LLM, be clear about what Jev actually does. The split card contrasts free generation — which writes token after token and “can return an unexpected fifth option” — with a typed decision constrained to refund, escalate, reply or ignore, where unlisted possibilities are discarded. That is the gap an Ollama-served chat model cannot close by prompting: asking a generative LLM for JSON gives you shaped text, not decision heads with calibrated probabilities behind each option. The interface may look similar; the guarantee is not there.

    Free generation versus typed decision split card where generation writes token after token but a typed decision must pick refund, escalate, reply or ignore and discards unlisted options
    Type safety means the model cannot return a fifth option.Watch at 2:22
  4. 4

    Run the local route: serve Kev-4B on 127.0.0.1

    The demonstrated local path is Kev — Jared Palmer’s Apache-2.0 decision adapters on Alibaba’s Qwen 3.5. The terminal serves the 4B model with one command: KEV_DTYPE=bf16 uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009. The application diff is a single line in the official Python SDK: address = TypeSafe cloud becomes address = http://127.0.0.1:8009. Same System One interface, only the address changes — the swap is the entire migration. This is the part people expect from “Ollama-style” local serving, except the model behind the port speaks decisions, not chat.

    Terminal serving Kev 4B locally with KEV_DTYPE bf16 uv run kev.serve on port 8009 while the SDK address swaps from TypeSafe cloud to http://127.0.0.1:8009
    Same System One interface, only the address changes.Watch at 5:22
  5. 5

    Measure the accuracy gap you accept for local

    The video’s comparison table prices the precision difference. On training-like tasks the local adapters sit at roughly 87% against hosted Jev’s ~88%. On the new-source test — material the models have not seen — the run shows 68.4% for Kev-0.8B, 83.7% for Kev-4B and 85.2% for Kev-9B, against 85.7% (dev) for hosted Jev. The trend line matters as much as the numbers: larger Kev models generalize better, and the smallest one gives back seventeen points on unseen sources. The 9B lands within half a point of the cloud; the 0.8B does not.

    New-source test table where Kev 0.8B scores 68.4 percent, Kev 4B 83.7 percent and Kev 9B 85.2 percent locally against hosted Jev at 85.7 percent
    On unseen sources the 9B lands within half a point of hosted Jev.Watch at 5:42
  6. 6

    Weigh latency: 329 ms to 2 seconds per decision on an M5

    Latency is where local models split into two different products. On an Apple M5, the video clocks 329 ms per decision for the 0.8B and 779 ms for the 4B — both faster than the hosted card’s 70–500 ms range on paper — while the 9B needs about 2,000 ms, four or more times the cloud round trip. Pair this with the accuracy table and the trade becomes explicit: the model that is fast enough to feel instant is the one that drops to 68.4% on unseen sources, and the model that matches the cloud needs a second of patience or a discrete GPU.

    Apple M5 latency chart giving 329 milliseconds per decision for Kev 0.8B, 779 milliseconds for 4B and roughly 2,000 milliseconds for 9B
    Small local models outrun the cloud chip; the accurate ones lag behind.Watch at 6:42
  7. 7

    Skip training: read scores from a model you already host

    SemIf (Theo Lee) is the zero-training entry in the field: instead of fine-tuning, it turns an unmodified open model into a decision engine by reading scores directly from a compressed Qwen 3.5 4B — the closest thing to “just use the model files you already downloaded.” In the video’s measurement it reaches 75.1% modal agreement with no fine-tuning, against hosted Jev’s published 78.5% on the same subset — 102 rows across 20 cases from TypeSafe’s selected set. That is roughly three points of agreement for zero training, on hardware as ordinary as the RTX 3090 class.

    SemIf local card recording 75.1 percent modal agreement from a compressed Qwen3.5 4B with no fine-tuning next to hosted Jev published results of 78.5 percent on the same 102-row subset
    Zero training, about three points of agreement — the cheapest local experiment.Watch at 7:42
  8. 8

    Survey the local field before you pick one

    No single project owns “local Jev.” The comparison table lines up four: Von, built on ModernBERT-Large, offers the broadest hardware reach and runs on a CPU with no GPU at all; Kev, on Qwen 3.5, carries the Jev-compatible interface; SemIf runs the technique on an unmodified open model you already host; OpenJev, on Qwen 3.5, reads images for visual judgment. None of them is Ollama — Ollama serves chat-style generation models, while these expose decision outputs. Pick by constraint: no GPU means Von, an existing SDK integration means Kev, zero training means SemIf, screenshots and receipts mean OpenJev.

    Open ecosystem table matching Von on ModernBERT-Large with no GPU needed, Kev on Qwen 3.5 with a Jev-compatible interface, SemIf on an unmodified open model and OpenJev for visual judgment
    Four open projects, four different reasons to reach for them.Watch at 11:12
  9. 9

    Decide when local is actually worth it

    The closing balance sheet settles the Jev vs local question. Local gains: no metered token charge, and context stays in your controlled environment. Your problems in exchange: hardware and electricity, deployment and monitoring, calibration and security. The video’s use-case cards mark the sweet spot as frequent, fixed answers that tolerate review — support routing, spam filtering, feedback tagging, lead scoring — and the case weakens when the output cannot be listed beforehand, on a spectrum from support queues (a few points acceptable) through credit approval to medical care. The final instruction applies to every model on this page: collect about 100 real examples from your own workload, measure accuracy, inspect the costly mistakes, and verify that a 90% confidence actually succeeds about nine times in ten before you automate anything.

    Balance sheet listing local gains such as no metered token charge and context staying in your controlled environment against hardware costs, deployment work and calibration duties
    Local wins on metering and data residency — and you inherit the operations.Watch at 12:42

Frequently asked questions

Can you run Jev itself in Ollama?

No. Jev is a hosted, closed-weight decision model, and Ollama serves chat-style open generation models — it has no equivalent of Jev’s typed decision heads or its calibrated probabilities. The video’s local alternatives do not run through Ollama either: Kev exposes a Jev-compatible System One interface from Qwen 3.5 adapters, SemIf reads scores directly from a model you host, and Von runs a small encoder on CPU. If what you want from “Ollama” is local control, those are the routes that actually preserve the typed-decision behavior.

Is a local decision model actually cheaper than hosted Jev?

Hosted Jev costs about $0.0004 per case at $0.042 per million input tokens, so at decision volumes that matter the meter is tiny. Going local removes the token charge but adds hardware and electricity, deployment and monitoring, and calibration and security work — the video’s balance sheet. Cost alone rarely justifies the move at this price point; privacy rules that forbid external processing are the reason that does.

How big is the accuracy and calibration gap for local models?

In the video’s run: 83.7% (Kev-4B) and 85.2% (Kev-9B) on the new-source test versus 85.7% (dev) for hosted Jev, with the 0.8B at 68.4%. Calibration improves with size — Brier score 0.460 for the 0.8B against 0.237 for the 9B on the test split, where lower is better. The zero-training SemIf route reaches 75.1% modal agreement versus 78.5% published Jev results on the same 102-row subset. Caveat from the video itself: most published benchmarks come from the projects themselves on different datasets, so treat them as indicators, not rankings.

When should I stay on hosted Jev instead of going local?

When privacy rules permit external processing, hosted wins on every axis the video measures: 70–500 ms responses, the best calibration, and no hardware to operate. The local case weakens when the output cannot be listed beforehand — the spectrum runs from support queues, where a few points of accuracy are acceptable, through credit approval, to medical care, where it is unacceptable. Frequent, fixed answers that tolerate review remain the sweet spot for local; everything else defaults back to the cloud.

Related guides

More video walkthroughs