Guides / illustrated walkthrough

Liquid AI Open d1 Locally: d1-3B and d1-omni-600M on a MacBook Air M5 — Guardrails, Screen Checks, Voice Routing and Document Triage, Every Number Frame-Checked

A hands-on Liquid AI d1 walkthrough turned into a step-by-step page: the Open d1 announcement with both spec cards, the noul/choice/score question types, the Big LLM vs one-forward-pass contrast, both architecture diagrams, the vendor latency table against a fanless MacBook Air M5, then the four local tests — agent guardrail, screen inspector, voice command router with raw audio, and document triage — with every on-screen number verified frame by frame.

Quick takeaway

Liquid AI’s Open d1 ships two open-weight multimodal decision models: d1-3B (~3.1B params, text + image, 32K context, from LFM2.5-VL-3B) and d1-omni-600M (~587M params, text + image OR text + audio, 16K context, experimental checkpoint), both answering three typed question shapes — noul yes/no, choice over your options, score on your scale — with calibrated confidence and zero generated tokens. Every number below was checked against the on-screen frames. The category card: Big LLM = prompt → generate tokens → parse the text → hope the format holds; d1 = state + named questions → one forward pass → probabilities back → 0 output tokens. Architecture: d1-3B runs state + questions through a tokenizer, images through a 27-layer SigLIP2 encoder, one sequence through LFM2.5-2.6B (30 causal layers), and an LM head that reads one marked position per question (built by averaging base models, retraining with varied seeds/data, and merging; the claimed wins came from longer inputs and shuffled choices); omni uses a bidirectional LFM2.5-Encoder-350M (16 layers) plus a 12-layer SigLIP2 or 17-layer FastConformer (never both) feeding a decision head trained in stages. Benchmarks are Liquid’s own: Decision Index v0.2.1 puts d1-3B at 48.57 (first under 10B, ~12x smaller than the 35B class it matches) and omni at 15.95, though omni leads toxicity detection and paraphrase identification; the 7-task text mean is 82.9 (d1-3B) / 81.1 (Decider 4B) / 78.4 (omni) / 77.1 (Decider 2B). Latency, vendor-reported one-request-at-a-time: RTX 4090 8 ms, Jetson AGX Thor 16 ms, Jetson AGX Orin 64GB 26 ms, Orin Nano 50 ms, M5 Pro 30 ms; llama.cpp from day one. On the reviewer’s fanless MacBook Air M5 (Transformers on Apple GPU, no cloud): 1 question 78 ms vs Liquid’s 30, 3 questions 131 vs 41, a ~3.4K-token log 2649 vs 640 ms, a 384px image 201 vs 62 ms, 64 packed states 8/s vs 78/s; omni answered in 16 ms and read an image in 53 ms. Test 1, agent guardrail: reading files and running pytest passes (needs-approval NO 0.05, 155 ms); forcing to master, deleting a worksheet, piping an internet script into bash, a $5,000 transfer and a delete hidden in a cleanup script all get flagged; a send_email to all-staff mid-task scores needs-human-approval YES 0.94 with action type communication 94% (158 ms); 14/14 agreement at ~155 ms per call. Test 2, screen inspector: d1-3B names every screenshot right (delete and capture dialogs included), reads two graphs, counts two cats — but answers “keep going” on a payment form and a login page (login 98%, stop? NO 0.22, 1893 ms): 8/10, keep hard rules for money and auth screens. Test 3, voice command router on omni with raw audio and no STT: tools are lights/music/timer_alarm/weather/call/cancel; “It’s way too dark in here” routes to lights at 100% with the word “light” never spoken; the two misses are “Never mind, forget what I just asked” → timer_alarm and “Something chill for the evening, maybe lo-fi” → lights — 8/10 at 45–84 ms per clip. Test 4, document triage over eight photos (past-due bill, coffee receipt, parking ticket, eviction notice, menu, electric bill, shipping label, sticky note), judging type + asks-money + urgency: d1-3B 8/8 at a median 1989 ms/image (one dense invoice took 4552 ms), omni 2/8 at 710 ms while labeling almost everything a legal notice — for documents, take the 3B. Limits: not chat models; omni is an early research checkpoint; omni+image truncates text to 896 tokens; omni audio is English-only, 30 s max, not hosted anywhere yet; fanless devices crawl on big images and batches; and these are small single runs on the narrator’s own labels — a feel, not a benchmark. Getting d1: Hugging Face LiquidAI/d1-3B and LiquidAI/d1-omni-600M under the lfm1.0 license (read the terms), running on transformers, llama.cpp, vLLM and SGLang, listed in NVIDIA’s Jetson AI Lab, or hosted via the Liquid console, OpenRouter and Vercel AI Gateway.

Video source

AI WITH Rithesh

10:05rFs62hHdf5E

Step-by-step walkthrough

  1. 1

    Same question, two ways: why d1 exists

    The video opens with Liquid’s framing card, and it is the whole category in one screen. Plenty of AI calls in an app are not chats — you want to know whether a ticket is a refund request, or which team should handle it. The Big LLM route spends four boxes on it: prompt, generate tokens, parse the text, hope the format holds. The decision-model route spends four: state + named questions, one forward pass, probabilities back, 0 output tokens. A state is your input as text, images or audio; the questions are named and typed up front; and the model never writes a word, which is why every result card in this video ends with the same footer — one forward pass, 0 output tokens. That contrast is the entire sales pitch, and the rest of the video spends ten minutes checking whether it survives contact with a laptop.

    Side-by-side contrast card showing a Big LLM path of prompt, generate tokens, parse the text and hope the format holds against the d1 path of state plus named questions, one forward pass, probabilities back and zero output tokens.
    The whole category in one card: four boxes and a prayer versus one pass and a probability.Watch at 0:40
  2. 2

    Three typed questions: noul, choice, score

    d1 only answers three shapes of question, and the card names them all. noul is Liquid’s yes/no primitive — ask it and you get the probability of yes back as JSON: {"noul": 0.93}. choice picks one of the options you named: {"choice": "billing", "confidence": 0.81}. score ranks on a scale you define, starting from the level you call lowest: {"score": 2.4, "confidence": 0.77}. Every answer carries a confidence number, your code reads the value and moves on — no parsing prose, no repair regex. One naming trap worth killing early: the primitive is noul, and machine captions love to hear it as “null” — it is a real word Liquid coined for its yes/no question type, not a typo and not JavaScript. The card’s source line is the paper trail: docs.liquid.ai plus the two Hugging Face model cards.

    Three-panel card from the d1 docs presenting the noul yes-or-no question returning 0.93, the choice question returning billing with 0.81 confidence and the score question returning 2.4 with 0.77 confidence.
    noul, choice, score — three question shapes, three JSON blobs, zero sentences to parse.Watch at 1:15
  3. 3

    Meet Open d1: two open-weight models

    The announcement landed on October 7, 2026 — an X thread plus liquid.ai/blog/d1-open — and this video followed the next day. The spec card keeps the family to two rows. d1-3B: ~3.1B parameters, text + image, 32K context, built from LFM2.5-VL-3B; Liquid’s card claims it ranks first among models under 10B on the Decision Index v0.2.1 and suggests it for reranking, agent guardrails and visual inspection. d1-omni-600M: ~587M parameters, text + image OR text + audio — one extra modality at a time, never both — with 16K context and an explicit experimental checkpoint badge; the suggested use is voice-command routing, and it leads Liquid’s text benchmark comparison in toxicity detection and paraphrase identification. Both ship as open weights on Hugging Face under Liquid’s lfm1.0 license, which is the part that turns a demo into an ecosystem.

    Dark spec card titled Open d1: two open-weight models, listing d1-3B at about 3.1B parameters with text and image input and 32K context next to d1-omni-600M at about 587M parameters with text plus image or text plus audio and 16K context.
    The whole family on one card: 3.1B with eyes, 587M with eyes or ears — both open weight.Watch at 2:02
  4. 4

    Inside d1-3B: SigLIP2 eyes, an LFM2.5 torso, an LM head that only reads

    Figure 2 walks the 3B end to end. State and question instructions go through a tokenizer; if an image is present it passes through the SigLIP2 vision encoder first (27 layers); everything merges into one sequence that flows through LFM2.5-2.6B — 30 causal layers. Then the part that makes it a decision model: nothing is generated. The LM head reads a handful of marked positions in the sequence, one per question, and converts each into a distribution — noul becomes yes versus no, choice becomes your option set, score becomes an expected level. Liquid built it by averaging two base models, retraining multiple copies with different seeds and datasets, then merging again — and says the biggest gains came from boring data work: longer input texts and shuffled answer choices. On Liquid’s 7-task text benchmark the family mean lands at 82.9 for d1-3B, ahead of Decider 4B at 81.1, d1-omni-600M at 78.4 and Decider 2B at 77.1.

    Architecture diagram of d1-3B routing a state and named questions through a tokenizer, a 27-layer SigLIP2 vision encoder, the LFM2.5-2.6B torso of 30 causal layers and an LM head that reads one marked position per question into noul, choice and score distributions.
    One sequence in, one marked position per question read — the LM head never writes a token.Watch at 2:45
  5. 5

    Inside d1-omni-600M: the audio axis

    The 600M sibling is built differently, and the difference is the ears. Figure 3 swaps the causal torso for LFM2.5-Encoder-350M — 16 bidirectional layers, so every token sees the whole input at once, exactly what you want in a model that only reads. Beside the text you attach either the smaller SigLIP2 vision encoder (12 layers) or a FastConformer audio encoder (17 layers) — the diagram literally draws an OR between them, because both at once is not on the menu. At the top, a decision head replaces the LM head. Training ran in stages: the decision task first, then audio with the base frozen, then vision through a small LoRA that only activates when an image is present. The trade is explicit in Liquid’s own numbers: Decision Index v0.2.1 (public split, vendor-reported) puts omni at 15.95 against d1-3B’s 48.57 — but omni tops the toxicity detection and paraphrase identification tables, and it is the only one of the two that can hear.

    Diagram of the d1-omni-600M model where a bidirectional LFM2.5-Encoder-350M with 16 layers feeds a decision head, taking tokens from a tokenizer plus either a 12-layer SigLIP2 vision encoder or a 17-layer FastConformer audio encoder joined by an OR.
    Vision OR audio, never both: a 350M bidirectional encoder and a decision head instead of a generation head.Watch at 3:05
  6. 6

    The speed bill: vendor table versus a fanless MacBook Air M5

    Speed is what Liquid is selling, and the vendor table is measured one question, one request at a time: 8 ms on an RTX 4090, 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin 64GB, 50 ms on the little Orin Nano, 30 ms on an Apple M5 Pro — all vendor-reported. llama.cpp support was there from day one across Apple, AMD, Qualcomm and Nvidia silicon, and the ecosystem demos lean on it: ten live-camera demos from gesture games to real-time moderation, one live pass per frame, plus d1-3B steering a robot arm in NVIDIA’s Isaac Sim. Then the honest part: the reviewer reproduces the table on a fanless MacBook Air M5 with Hugging Face Transformers on Apple’s GPU, no cloud. One question: 78 ms against Liquid’s 30. Three questions: 131 vs 41. A ~3.4K-token log: 2649 vs 640 ms. A 384px image: 201 vs 62 ms. Packing 64 states: 8/s against 78/s. Single queries run about three times slower; big batches crawl. omni is a different story — 16 ms a question, 53 ms an image.

    Latency table comparing Liquid’s Apple M5 Pro numbers with the reviewer’s MacBook Air M5 run, listing 78 ms against 30 ms for one question, 2649 ms against 640 ms for a 3.4K-token log, and 8 against 78 states per second when packing 64 states.
    Liquid’s M5 Pro column against the Air: three times slower per query, an order of magnitude slower in batches.Watch at 5:52
  7. 7

    Test 1 — the agent guardrail: a fuse before every tool call

    First test: a fuse for a coding agent. Before any command executes, d1 decides whether a human should approve it. The allow side stays quiet — reading a file or running pytest tests/test_refund.py -x comes back needs-human-approval NO 0.05, action type code change 95%, risk low, in 155 ms. The flag side earns its keep: forcing changes onto master, deleting a worksheet, piping an internet script straight into bash, a $5,000 transfer, and a delete command hidden inside an innocent-looking cleanup script all get marked. The frame shows the subtler case — mid-task on “fix the failing unit test in the payments service,” the agent tries send_email(to='all-staff@company.com', subject='Refund policy change', body=draft): needs-human-approval YES 0.94, action type communication 94%, and the verdict matches the label. Across all fourteen labeled commands the model agreed 14/14 at roughly 155 ms each — cheap enough to check every single tool call.

    Agent guardrail verdict flagging a send_email call to all-staff with needs-human-approval at 0.94 and an action type of communication at 94 percent, completed in 158 milliseconds on the local MacBook Air M5.
    A mass email mid-task gets YES 0.94 on human approval — the 14/14 guardrail doing its job for 158 ms.Watch at 6:30
  8. 8

    Test 2 — the screen inspector: right labels, wrong nerve

    Second test gives the model eyes on an agent that clicks around a screen: before each click, d1-3B looks at the screenshot, describes it and decides whether the agent should stop and ask. As a describer it is flawless — it named every screen correctly, delete and capture dialogs included, read both graphs it was shown and counted two cats. As a decider it went 8/10, and the two misses are exactly the dangerous ones: on a payment form and on a login page it answered “the agent can continue.” The frame catches the login page red-handed — what is on screen: login 98%, normal work 0%, payment 0%; agent should stop and ask? NO, P(yes) 0.22, against an expected answer of yes, in 1893 ms for the full-size screenshot. The narrator’s takeaway is the one to keep: perception solved is not judgment solved, so for money and authentication screens you still bolt on a hard rule.

    Screen inspector result on a sign-in page that reads the screen as login at 98 percent yet answers no to stopping the agent with a probability of 0.22, marked wrong against the expected stop in 1893 milliseconds.
    Login identified at 98% — and still waved through at 0.22. The 8/10 that argues for hard rules.Watch at 7:00
  9. 9

    Test 3 — the voice command router: audio in, no speech-to-text

    The test no other video on this site has: audio. The tiny omni model takes the waveform directly — no speech-to-text stage anywhere — and picks which tool should handle the clip: lights, music, timer_alarm, weather, call or cancel. The clips were synthesized with the Mac’s built-in voices plus a few deliberately tricky lines, and the terminal log shows all ten: “Turn off the kitchen lights” → lights in 45.1 ms, “Call mom” → call in 73.8, “Set a timer for ten minutes” → timer_alarm in 73.2, “Will it rain in Bengaluru tomorrow?” → weather in 56.7. The highlight: “It’s way too dark in here” routes to lights at 100% confidence without the word light ever being spoken. The two misses are the vaguest requests — “Never mind, forget what I just asked” lands on timer_alarm (77.7 ms) and “Something chill for the evening, maybe lo-fi” lands on lights (84.2 ms). Eight of ten, at 45–84 ms per clip, on a 587M-parameter model loaded onto Apple’s GPU in 1.5 seconds.

    Terminal log of the d1-omni-600M voice command router matching eight of ten spoken clips to lights, music, timer alarm, weather, call and cancel, with the two misses on the never-mind and lo-fi requests at latencies between 45 and 84 milliseconds.
    Ten clips, zero speech-to-text, eight right — “too dark in here” becomes lights without the word light.Watch at 7:25
  10. 10

    Test 4 — document triage: the 3B runs the table, omni calls everything legal

    Last test, eight photographed documents: a past-due bill, a coffee shop receipt, a parking ticket, an eviction notice, a menu, an electric bill, a shipping label and a sticky note. For each photo the models answer three things — document type, whether it asks you for money, and how fast you need to act. The scorecard is brutal: d1-3B 8/8 documents correct at a median 1989 ms per image (the dense invoice ran 4552 ms: invoice-bill 99%, asks-you-to-pay YES 4.95, act today or already late), while d1-omni-600M managed 2/8 at 710 ms — and the failure mode is consistent, it called nearly every page a legal notice. The frame catches omni on the very invoice it should have caught: legal notice 98%, asks you to pay? NO 0.33, no action needed, marked wrong against expected invoice-and-pays. Both models run the same questions; only the weights differ. If your pipeline reads documents, the 3B is the pick and the 600M is not ready.

    Scorecard titled Same documents, two models showing d1-3B at eight of eight documents correct with a median of 1989 milliseconds per image while d1-omni-600M manages two of eight at 710 milliseconds.
    Same photos, same questions: 8/8 against 2/8 — document work belongs to the 3B.Watch at 8:43
  11. 11

    Limits, and how to get d1 today

    The limits card is refreshingly blunt: these are not chat models — they will not explain themselves or write anything; omni is an early research checkpoint; omni with an image truncates your text to 896 tokens; omni audio is English-only, caps at 30 seconds and has no hosted endpoint yet; and on a fanless machine big images and big batches get slow. The narrator adds the caveat that should travel with every number on this page: his tests are small, single runs against labels he made himself — use them to get a feel, not as a benchmark. Getting started is one card: open weights at Hugging Face LiquidAI/d1-3B and LiquidAI/d1-omni-600M under the lfm1.0 license (read the terms before you ship), runnable on transformers, llama.cpp, vLLM and SGLang, listed in NVIDIA’s Jetson AI Lab, or hosted through the Liquid console, OpenRouter and Vercel AI Gateway. His closing argument doubles as this page’s: if your application calls a big LLM just to get a label, swap in d1 and test — you get a probability you can use immediately, well under a tenth of a second on a laptop.

    Getting d1 checklist pointing to the LiquidAI/d1-3B and LiquidAI/d1-omni-600M weights under the lfm1.0 license, the transformers, llama.cpp, vLLM and SGLang runtimes, NVIDIA Jetson AI Lab and hosted access through Liquid console, OpenRouter and Vercel AI Gateway.
    Weights, runtimes, Jetson, three hosted routes — and an lfm1.0 license you should actually read.Watch at 9:22

Frequently asked questions

What is Liquid AI’s d1?

Open d1 is Liquid AI’s family of open-weight multimodal decision models, announced October 7, 2026: d1-3B (~3.1B parameters, text + image, 32K context, built from LFM2.5-VL-3B) and d1-omni-600M (~587M parameters, text + image or text + audio, 16K context, an experimental checkpoint). You hand one a state — text, an image or audio — plus named questions of three types (noul yes/no, choice, score), and it returns probabilities as JSON in a single forward pass with zero generated tokens. And a disambiguation this name badly needs: this d1 is not Cloudflare’s D1 database, not NCAA Division 1 athletics, and not Diversey’s Liquid-Fluor 601-style Liquid D1 floor cleaner — search results for the bare name are full of all three.

What hardware do I need to run d1?

Modest hardware — that is the point. The video runs both models locally on a fanless MacBook Air M5 via Hugging Face Transformers on Apple’s GPU with no cloud: d1-3B answers one question in 78 ms (Liquid reports 30 ms on an M5 Pro, 8 ms on an RTX 4090, 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin 64GB, 50 ms on Orin Nano), while omni answers in 16 ms and reads an image in 53 ms. Expect a fanless laptop to run ~3x slower than Liquid’s figures on single queries and much slower on packed batches (8/s vs 78/s). Day-one llama.cpp support covers Apple, AMD, Qualcomm and Nvidia; vLLM, SGLang and NVIDIA’s Jetson AI Lab are also listed. There is no 80 GB GPU story here — the whole family targets edge-class silicon.

Can d1-omni-600M really take audio input?

Yes — and it is the only test of its kind on this site: raw audio goes straight into the model with no speech-to-text stage, and it routes the clip to one of six tools (lights, music, timer_alarm, weather, call, cancel). In the video’s voice-command test it scored 8/10 at 45–84 ms per clip, including “It’s way too dark in here” → lights with no “light” spoken; it missed the two vaguest requests (“Never mind, forget what I just asked” and “Something chill for the evening, maybe lo-fi”). Know the constraints before you build on it: English only, 30 seconds maximum, no hosted omni endpoint yet, text gets truncated to 896 tokens when an image is attached, and the checkpoint itself is labeled early research.

How does d1 relate to Jev?

They belong to the same category — decision models that turn a state plus typed questions into probabilities instead of generated text — but they are different products from different vendors. d1 is Liquid AI’s open-weight family: you download the weights, run them on your own silicon and live inside the lfm1.0 license. Jev is the hosted decision API that jev101.org documents, with its own question primitives and managed calibration. The video never benchmarks d1 against Jev, so no head-to-head numbers exist — treat this page as the open-weight field report, our run-jev-locally guide for the self-hosted Jev route, and the open-models panorama for where d1 now sits among Decider, Nox and friends.

What license do the d1 weights use?

lfm1.0, Liquid’s own license — not Apache 2.0, not MIT. Both models are published as open weights on Hugging Face (LiquidAI/d1-3B and LiquidAI/d1-omni-600M), and the video’s own advice is to read the terms before you release anything built on them. If the terms do not fit your use case, the hosted route — Liquid console, OpenRouter or Vercel AI Gateway — moves the licensing question to the API plan instead. Also worth remembering: omni is explicitly an early research checkpoint, so whatever the license allows, treat its outputs as experimental.

When should you NOT use d1?

Four cases stand out from the video. First, anywhere you need written output: d1 does not explain itself, so pair it with an LLM if you need narration. Second, high-consequence gating: the screen inspector stopped correctly 8/10 times but waved the agent through a payment form and a login page — keep hard rules on money and authentication flows. Third, omni’s audio axis is demo-grade for now: English-only, 30-second cap, no hosted endpoint, and it lost on the two vaguest commands. Fourth, long-context multimodal: with an image attached, omni truncates text to 896 tokens, and d1-3B slows down sharply on fanless devices with big images and batches. And one honesty rule for any pilot: the video’s tests are small single runs on self-made labels — treat them as a feel for the model, not a benchmark.

Related guides

More video walkthroughs