Guides / illustrated walkthrough

Ollama Decision Models: Run tev1 and Nimble Locally (Tested on an 8 GB Card)

A hands-on test of Ollama 0.35’s new /v1/systemone endpoint: pull tev1 (4B/0.8B, Together AI) and Nimble (9B, Bespoke Labs), route tickets and moderate messages in a few hundred milliseconds, then compare them against a JSON-schema-forced chat model on the same 30 tickets.

Quick takeaway

Ollama 0.35 (released alongside the 2026-09-29 announcement) added a /v1/systemone endpoint — not a chat endpoint — that runs Jev-style decision models locally. Two families are live: tev1 from Together AI (4B and 0.8B, further-trained Qwen 3.5) and Nimble 9B from Bespoke Labs (Apache 2.0). The video walks the whole path on an 8 GB laptop: pulling the models, learning the finicky request shape through three validation errors, and running a four-tab Decision Lab (triage, moderation, router, grader). On the creator’s own 30-ticket test — a good start, not a benchmark — tev1 4B scored 100% on team routing at 470 ms using 4.7 GB, Nimble hit 96.7% at 1.5 s but spills onto the CPU, and the 0.8B runs in under a gigabyte. The most instructive part: a plain Qwen 3.5 chat model forced into a JSON schema matched the accuracy and beat the speed (268 ms) — what it cannot give you is calibrated confidence (every decision-model error sat below 0.9) or consistency (the same ticket came back “billing” 7 times out of 8). That is the whole local-decision-model pitch in one experiment.

Video source

Prompt Engineer 48

8:18sc-qe-gNwKk

Step-by-step walkthrough

  1. 1

    Start from the announcement: Ollama now speaks System One

    The video opens on Ollama’s blog post dated September 29, 2026: “Ollama now supports Jev-style decision models.” Decision models based on TypeSafe’s Jev API now run through a new endpoint on your own machine — the post calls out no additional costs, lower latency than a hosted call, and three new decision models available the same day. The fine print that matters later: the new API requires Ollama 0.35 or later, so check ollama --version before anything else. The creator had exactly 0.35.0 — the minimum.

    Ollama blog post from September 29 2026 announcing Jev-style decision models with no additional costs, lower latency locally, and three new models via Ollama
    The 2026-09-29 announcement: Jev-style decisions, local, no extra cost.Watch at 0:25
  2. 2

    Meet the cast: tev1 in two sizes and Nimble 9B

    Three models land in ollama list. tev1 comes from Together AI in a 4B (“a 4B decision model from Together AI for fast classification,” a 4.5 GB download) and an 0.8B (811 MB) — both further-trained on Qwen 3.5. Nimble is a 9B from Bespoke Labs under Apache 2.0 (a 7.5 GB file). TypeSafe publishes benchmarks of 73.3% for tev1 4B and 75.7% for Nimble — the video is careful to flag these as the vendors’ numbers, not his, and then runs its own test. Both model cards are one ollama run away.

    ollama list output showing tev1 0.8b at 811 MB, nimble latest at 7.5 GB, and tev1 latest at 4.5 GB on a local machine
    811 MB, 4.5 GB, 7.5 GB — the whole decision-model menu fits a laptop.Watch at 2:13
  3. 3

    Learn the request shape through three validation errors

    This is not a chat endpoint — it is /v1/systemone on port 11434, and the request takes a model, a state, and named questions. The video’s first curl fails: “question team: instructions must be a nonempty string, object, or array.” The second try fails nicer: choice criteria must associate option keys with descriptions (a map like billing → payments and returns). The third error teaches that a score question’s criteria is an array of descriptions, not a map. Every message tells you exactly what is missing — but the lesson stands: rigid schemas, strict validation. Copy the shape before you freestyle.

    Terminal showing curl to localhost 11434 v1 systemone returning the error question team instructions must be a nonempty string object or array
    Error one of three: every question needs an instructions string.Watch at 2:32
  4. 4

    One call, three typed questions, four output tokens

    Once the shape is right, a single request asks three questions about the same ticket and gets everything back at once: billing at 0.9975 (choice), customer-anger at 0.87 (noul), urgency at 0.65 (score). The counter that explains the speed: 4 output tokens. Nothing is generated — the answers and probabilities arrive in one forward pass, which is why a decision endpoint can sit on the hot path of every message. Up to 64 questions can ride in one call.

    Result cards showing billing 0.9975, angry 0.87, urgency 0.65 and 4 output tokens from one systemone request
    0.9975 billing + 0.87 angry + 0.65 urgency — and only 4 tokens out.Watch at 3:15
  5. 5

    Watch it route: the four-tab Decision Lab app

    The creator wraps the endpoint in a small Gradio app. On the triage tab, “My card was charged twice and I want a refund now!!” routes to billing at 0.99 confidence in 733 ms (794 tokens in, 4 out). An “API returns 500, production is down” message goes to tech at 0.94. The persuasive one: “Do you have a discount for non-profits?” lands on sales — no billing keyword anywhere. The model reads intent, not vocabulary, and every answer carries a number you can gate on.

    Decision lab triage tab routing a double-charge refund ticket to billing with 0.99 confidence, angry 0.88, urgency 0.60 at 733 ms
    Refund ticket → billing 0.99 in 733 ms; intent, not keywords.Watch at 3:35
  6. 6

    Moderation is three yes/no gates in one pass

    The moderation tab fires three noul questions per message — spam, personal data, toxicity. A fraud message (“Congrats!! You won 1000 USD. Click http://free-cash.xyz and send your card number”) comes back spam 0.97 BLOCK, pii 0.95 BLOCK, toxic 0.09 pass. The app’s rule is simple: block anything over 0.75. Each gate exposes its own probability, so your policy can treat a card-number leak differently from a scam link without a second model call.

    Moderation tab showing a scam lottery message blocked with spam 0.97 and pii 0.95 while toxic scores 0.09 under a 75 percent block rule
    Spam 0.97, pii 0.95, toxic 0.09 — three gates, one forward pass.Watch at 4:10
  7. 7

    The honest moment: a friendly message scores spam 0.66

    “Hi, I am John, call me on +1 415 555 0134” — personal data fires correctly at 0.98, but tev1 rates spam at 0.66 and Nimble at 0.49. A phone-number intro is unusual, not fraudulent, and the models wobble. That wobble is why the creator sets the block threshold at 0.8 instead of the default-ish 0.5: thresholds are a policy decision you make after watching real false positives, not a number you copy. This one minute is the most transferable lesson in the video.

    Moderation checks for a friendly message showing pii 0.98 blocked while spam only scores 0.66 illustrating a false positive near the threshold
    A polite message at spam 0.66 — why his block line sits at 0.8.Watch at 4:22
  8. 8

    The 30-ticket self-test: tev1 4B never misses a route

    He writes 30 support tickets himself, each labeled with the correct team and anger level — explicitly a personal test, not a public benchmark. Results: tev1 0.8B gets 86.7% on team and 76.7% on anger at 235 ms in 0.9 GB; tev1 4B gets a perfect 100% on team and 83.3% on anger at 470 ms in 4.7 GB fully on the GPU; Nimble 9B leads anger detection at 96.7% and matches 96.7% on team — but takes 1.5 s a call. His verdict for an 8 GB laptop: tev1 4B is the one to run; Nimble needs 12 GB or more to shine; the 0.8B is the when-memory-is-tight fallback.

    Bar chart of team accuracy on 30 self-labeled tickets with tev1 0.8b at 86.7 percent, nimble 9b at 96.7 percent and tev1 4b at 100 percent
    His 30 tickets: 0.8B 86.7%, Nimble 96.7%, tev1 4B a clean 100%.Watch at 6:12
  9. 9

    Why Nimble lags: 10 GB does not fit an 8 GB card

    ollama ps explains Nimble’s 1.5 s: the loaded model wants about 10 GB against an 8 GB card, so roughly 40% of the work runs on the CPU. There is also a concurrency trap — calling Nimble while tev1 was still loaded returned a host-side 500 (CUDA host-buffer allocation failed) until he ran ollama stop tev1. On small cards, one decision model at a time is the operating rule; plan your unloads like you plan your deployments.

    ollama ps output showing nimble using 10 GB with 40 percent CPU and 60 percent GPU and a note that 10 GB does not fit an 8 GB card
    Nimble at 10 GB: 40% CPU spillover is the price of the accuracy lead.Watch at 6:06
  10. 10

    The control experiment: force a chat model into JSON

    “Why not just ask a regular model to output JSON?” He pulls Qwen 3.5 4B, forces a JSON schema, and runs the same 30 tickets: 96.7% on team, 96.7% on anger, at 268 ms — accuracy tied with Nimble and speed ahead of everything. So what does the decision model buy you? Two things the chat path cannot ship. First, calibrated probability: tev1 4B made zero routing errors, the 0.8B’s four errors all sat below 0.77, and Nimble’s only error scored 0.83 — “below 0.9, send it to a person” is a rule you can actually run. Second, consistency: the chat model answered the same ticket “billing” seven times out of eight and “tech support” once, and reports nothing about its own confidence.

    Timing cards comparing 268 ms for JSON-schema-forced Qwen 3.5 4B chat against 470 ms tev1 4B and 1.5 s nimble 9B with the note accuracy is a tie speed chat wins
    Accuracy: a tie. Speed: chat wins. Confidence and consistency: Jev-style wins.Watch at 6:52
  11. 11

    Confidence against ground truth: the gate becomes real

    The closing chart plots every answer from the 30-ticket run as confidence versus right-or-wrong. The pattern that makes automation safe: tev1 4B’s wrong answers simply do not exist, Nimble’s single miss sits at 0.83, and the 0.8B’s mistakes cluster under 0.77. With errors living below the line, a threshold like 0.9 routes the doubtful cases to a human and almost never gives up a correct answer — the same confidence-gated pattern the hosted Jev API sells, now running next to your app.

    Scatter plot of team answers on 30 tickets showing confidence versus correct or wrong with red wrong answers clustered under the 0.9 send-to-human line
    Wrong answers live below 0.9 — that is what makes the threshold honest.Watch at 7:00
  12. 12

    The catch list: experimental labels and a 30-ticket horizon

    The verdict slide splits cleanly. The good: local, free, no API key, answers in hundreds of milliseconds, probabilities you can act on, 64 questions per call. The catch: the request shape is picky, Nimble will not fit an 8 GB card, and the tev1 models are marked experimental — the official page says they can be wrong. His own sign-off is the right model for yours: 30 tickets is a good start, not a proof. Validate on your own labeled data before any auto-action, and keep a human review path under the threshold.

    Warning slide reading do not let it be the only check with notes that the official page says it can be wrong and confidence under 0.9 goes to a human
    A good start, not a proof — the experimental label is doing real work.Watch at 7:38

Frequently asked questions

What are the Ollama decision models tev1 and Nimble?

Two Jev-style decision-model families available through Ollama 0.35’s /v1/systemone endpoint. tev1 comes from Together AI in 4B (4.5 GB) and 0.8B (811 MB) sizes, both further-trained on Qwen 3.5; Nimble is a 9B from Bespoke Labs under Apache 2.0 (7.5 GB download, about 10 GB loaded). TypeSafe publishes 73.3% (tev1 4B) and 75.7% (Nimble) on their public benchmark; the video’s own 30-ticket test put tev1 4B at 100% team-routing accuracy and Nimble at 96.7%.

How do I call the Ollama System One endpoint correctly?

POST to http://localhost:11434/v1/systemone with a JSON body containing model, state (your text), and questions. Three validation rules the video learned the hard way: every question needs a nonempty instructions string; a choice question’s criteria is a map of option keys to descriptions; a score question’s criteria is an array of descriptions. Up to 64 questions can share one request, and answers come back with per-option probabilities in a single pass.

Which model should I run on an 8 GB GPU?

tev1 4B. It went 100% on the video’s 30-ticket routing test at 470 ms while fitting entirely in 4.7 GB of VRAM. Nimble 9B is more accurate on anger detection (96.7%) but loads about 10 GB, spilling ~40% of compute onto the CPU at 1.5 s per call — give it 12 GB or more. The 0.8B (811 MB, 235 ms) is the fallback for tight-memory machines, at a real accuracy cost (86.7%). Also: unloading one big model before loading another avoids host-side CUDA 500s.

A JSON-forced chat model matched the accuracy — why bother with a decision model?

Because accuracy is not the whole contract. In the video, Qwen 3.5 4B with a forced JSON schema scored the same 96.7/96.7 at a faster 268 ms. What it could not provide: calibrated confidence (every decision-model error sat below 0.9, so a threshold can catch it — the chat model offers nothing similar) and consistency (it answered the same ticket “billing” 7 of 8 times and “tech support” once). If you only need an answer and can tolerate silent flips, schema-forced generation is fine; if you need to automate on a number, you need calibrated probabilities. The trade-off is compared in detail in our Jev vs JSON mode recipe.

Are these models production-ready?

Treat them as experimental. The official tev1 page states the models can be wrong, and the video closes with “my test was 30 tickets — a good start, not a proof.” The honest pattern from the video: set thresholds from your own false positives (the friendly message that scored spam 0.66 is the canonical example), route anything under ~0.9 to a human, and validate on labeled data from your own workload before any auto-action.

How is this different from the hosted Jev API?

Same interaction shape — state plus typed questions in, answers plus probabilities out — but the compute is yours: no API key, no per-token bill, no data leaving the machine, at hundreds-of-milliseconds latency on consumer hardware. The hosted Jev API remains the zero-ops option with the calibrated RLCD probabilities and higher accuracy ceiling; the local route trades some accuracy for cost, privacy, and latency control. The cloud-versus-local decision is covered in our Jev vs Ollama guide.

Related guides

More video walkthroughs