Guides / illustrated walkthrough
PPLX Decider Tutorial: Perplexity's Open 27B Decision Model, Its Decisions API, and a 12-Ticket Triage Run
Codedigipt's walkthrough of PPLX Decider v1 turned into a step-by-step page: the Apache 2.0 weights on Hugging Face, the 11-benchmark panel against Jev, the Decision Index scores, the three question primitives behind the Decisions API, image-carrying state, and the official triage cookbook that turns 12 tickets into 3 escalations — with the pricing change and the v1.1 update noted where the video is now out of date.
Quick takeaway
This is Codedigipt's seven-and-a-half-minute tour of PPLX Decider, Perplexity's open-source alternative to Jev in the decision-model class — video ID kDyTyzZWg9k, published October 2, 2026. The model lives at huggingface.co/perplexity-ai/pplx-decider-v1-27b: Apache 2.0 license, Text Classification pipeline, multimodal and custom-code tags, fine-tuned from Qwen3.8-27B per the card's own metadata (the architecture tag reads qwen3_5). The video's evidence, in order: an 11-benchmark accuracy panel on a fixed 2,710-row set, higher is better, September 2026, where Decider takes FinancialPhraseBank (84.18% vs Jev's 76.98%) and RAGTruth (88.80% vs 77.27%) while Jev holds WinoGrande (90.70% vs 83.30%), BBH (94.27% vs 82.80%), JudgeBench (78.57% vs 78.29%) and JevBench public hard (73.27% vs 70.30%) — and the Overall row lands at 85.71% vs 84.51% vs the base model's 74.76%, which is the '85.71%' the narrator cites; the card keeps the honest line 'These are Perplexity-reported results; they do not establish that the model will outperform every real-world decision task.' On the community Jev Decision Index (70 open reproductions, 120K decisions per model across 43 benchmarks, suite v0.2.1 updated 2026-09-28), the board's own hover cards read Jev 57.9, rank #1, 524 ms median, 'hosted API - HTTP round trip', versus pplx-decider-v1-27b 56.4, rank #4, 101 ms median, size 28B, full fine-tune — the narrator says '57.1' and '1.2 points', but the screen says 57.9, so quote the card (the gap is 1.5). Timeliness, verified October 7, 2026: the Decisions API input price the video quotes at $0.04 per million tokens has been halved to $0.02/M (output tokens free, no per-request fee — same figure our comparison recipe already uses), the docs now default to pplx-decider-v1.1-27b (v1 still supported) whose card reports a Decision Index of 61.56 vs Jev's 57.9 — the lead has flipped — and the endpoint is POST https://api.perplexity.ai/v1/decisions with Bearer auth. The API contract, as shown: state (text, JSON, or images) plus named questions; noul returns the probability of yes from 0 to 1 (criteria optional, nothing at all returns 400); choice returns the most likely option, probabilities for all 1 to 255 options, and confidence; score returns the probability-weighted level index on a 2-to-10-level rubric plus legend and full distribution; a request takes 1 to 128 questions under roughly 262,144 input tokens; images ride in state as OpenAI-style base64 data URLs (PNG/JPEG/WebP only — the API never fetches URLs), read in 32×32-pixel tiles with a 2,048-tile ceiling (~1440×1440; oversized images stall about a minute then 504), at roughly 1,000 input tokens per megapixel; confidence is the model's own certainty, not the top probability, and identical requests can differ in the second decimal — set thresholds with margin. The payoff is Perplexity's official triage cookbook: three probabilistic questions per ticket (needs human? intent? severity?), one request per ticket, routes to escalate / review / queue / auto_reply / close; the terminal run classifies T-1001 through T-1012 in about 2.5 seconds (most requests 150–250 ms at 590–659 input tokens), escalates exactly three (T-1001 outage, P(crit)=0.97; T-1006 security, 0.99; T-1007 outage, 0.85), and prints 'Add --draft-escalations to send the escalated tickets to the Agent API' — the cheap model decides, the expensive model sees only what deserves it. Self-hosting is the other road: Python 3.12+ and a CUDA GPU with room for about 49 GiB of weights plus working memory.
Video source
Codedigipt
Step-by-step walkthrough
- 1
The release: Jev's decision behavior, open weights
The video opens on the Hugging Face model card for perplexity-ai/pplx-decider-v1-27b — Perplexity's own org. The tag row reads Text Classification, PyTorch, Safetensors, qwen3_5, classification, multimodal and custom-code, and the license chip says apache-2.0. The card's own metadata is the cleanest definition you will get: a decision model fine-tuned from Qwen3.8-27B. The narrator's framing is the news: Jev is the closed decision model everyone benchmarks against, and Perplexity just published an open-weights cousin — which means you can inspect it, self-host it (the card later admits that costs about 49 GiB of GPU memory), or fine-tune it further. Two things have changed since the video: the card's download counter keeps climbing, and a v1.1 checkpoint (pplx-decider-v1.1-27b, published October 5, 2026) now exists — today's API defaults to it, with v1 still selectable.

The card the video opens on: Perplexity’s own org, Apache 2.0, a fine-tune of Qwen3.8-27B — not a rerouted API wrapper.Watch at 0:10 - 2
The vendor panel: 11 benchmarks, September 2026
The video scrolls to the numbers. The table is headed 'Accuracy on a fixed 2,710-row panel. Higher is better. September 2026.' with three columns: Jev, Qwen3.8-27B, pplx-decider-v1-27b. Read it as a mixed result, not a sweep: the Decider takes FinancialPhraseBank outright (84.18% vs 76.98%) and RAGTruth by a wide margin (88.80% vs 77.27%), while Jev holds WinoGrande (90.70% vs 83.30%), BBH (94.27% vs 82.80%), JudgeBench (78.57% vs 78.29%) and JevBench public hard (73.27% vs 70.30%). The most consistent pattern is the one the narrator emphasizes: the fine-tune beats its own Qwen3.8-27B base nearly everywhere — decision-specific training is doing real work. The on-screen figures match the model card's benchmark table on Hugging Face row for row.

Not a sweep: Decider takes FinancialPhraseBank and RAGTruth; Jev keeps WinoGrande, BBH and JudgeBench.Watch at 0:30 - 3
The Decision Index: 56.4 against Jev's 57.9
Next the video pulls up the community scoreboard: the multimodalart Jev Decision Index — 70 open reproductions of TypeSafe Jev's decision model, 120K decisions per model across 43 benchmarks, suite v0.2.1, updated 2026-09-28. On the summary bars Jev sits first and pplx-decider-v1-27b just below. Quote from the board, not the voiceover: the narrator says '57.1' and 'only 1.2 points', but the page's own hover cards read Jev 57.9 (rank #1, 524 ms median latency, 'hosted API - HTTP round trip, not a local GPU run') versus pplx-decider-v1-27b 56.4 (rank #4, 101 ms median, size 28B, technique full fine-tune, fine-tuned from Qwen/Qwen3.8-27B) — a gap of 1.5. The latency pairing is the interesting part: the closed model answers over the network while the open one runs adjacent to your code. And the standings have already moved again: Perplexity's v1.1 card (October 2026) reports a Decision Index of 61.56 against Jev's 57.9 — 'outperforming Jev by more than 3.5 points' — on the same backbone.

Hover before you quote: the board’s own cards say 56.4 at #4 (101 ms) and Jev 57.9 at #1 (524 ms) — not the narrator’s 57.1.Watch at 1:16 - 4
The hosted route: the Decisions API
From weights to the managed product. The docs landing defines it in one line: the Decisions API answers questions about your content with probabilities instead of text — you send the content as state (text, JSON, or images), attach one or more named questions, and get one answer per question. Pricing is where the video has aged: the narrator quotes $0.04 per million input tokens with output completely free — the logic being that answers are probabilities and option names, so there is almost nothing to bill on the way out. Perplexity has since halved input pricing to $0.02 per million tokens (verified on the official pricing page, October 7, 2026 — the same figure our Jev-vs-Perplexity recipe already uses); output tokens are still free and there is no per-request fee. One more drift to know if you follow the video's screenshots: today's quickstart sends POST https://api.perplexity.ai/v1/decisions with Bearer authentication.

State in, typed probabilities out — and since October 2026, $0.02/M input, not the $0.04 the video quotes.Watch at 1:45 - 5
noul: the yes-or-no primitive
The first of the three question types. noul takes a yes/no question or a statement to check and returns the probability of yes, from 0 to 1. You can attach criteria that defines what counts as yes and what counts as no; send a noul with neither instructions nor criteria and the API returns a 400. (The video's auto-transcribed captions garble the name into 'null' — the docs are unambiguous.) The frame's example asks 'Does the review report a product defect?' with true and false criteria. In practice this is your gate primitive: compare the returned probability against a cutoff you own, and treat values near 0.5 as the model telling you it is unsure.

One number out: the probability of yes — gate it against a threshold you own.Watch at 2:25 - 6
choice: routing across up to 255 options
choice is the classification primitive: pick one of the options you define. criteria maps each option name to a description of when it applies — pass null as a description to let the name speak for itself — and a question accepts 1 to 255 options. The answer carries choice (the option with the highest probability), probabilities covering every option (summing to about 1), and a confidence value. The narrator's example routes incoming messages into billing, technical support or sales; the JSON on screen routes a support ticket across critical, support, escalations, billing and other. This is the slot where you would otherwise prompt a chat model for a label and write a parser for the reply — here the distribution arrives pre-parsed.

The routing primitive: named options, per-option probabilities, one winner — no reply-parsing layer.Watch at 2:45 - 7
score: grading on an ordered rubric
Third primitive: score rates the content on an ordered rubric. criteria is an ordered array of level descriptions — the array index is the level's score, starting at 0 — with up to 10 levels and at least two (a single level has nothing to decide, so it always returns 0 with probability 1). The answer has score, the probability-weighted average of the level indices, which can fall between two levels; legend, mapping each index back to your rubric entry; probabilities for every level; and confidence. The frame's example grades 'How severe is the reported problem?' across Cosmetic, Inconvenient, Product unusable — exactly the dial the triage cookbook leans on later. Two calibration notes from the current docs: confidence is the model's own certainty estimate, not the top probability (it drops when the runner-up is close), and identical requests can differ in the second decimal place, so set thresholds with margin.

The severity dial: an expected score that lands between levels, with the full distribution attached.Watch at 3:05 - 8
Screenshots are first-class state
state is not limited to text: pass it as an array and each image rides in an OpenAI-style image part containing a base64 data URL — the video shows the request card with 'model': 'pplx-decider-v1-27b' and an image_url part asking which color the square is. The rules worth memorizing: PNG, JPEG and WebP data URLs only, and the API never fetches a URL — an http or https image URL returns 400. Images are read in 32×32-pixel tiles; keep each image at or under 2,048 tiles (1440×1440 and 2048×1024 fit; 1600×1310 does not), and an oversized image does not error — it waits about a minute and then returns 504, so resize before you send. Image tokens count toward input: about 1,000 tokens per megapixel in Perplexity's tests. The official use-case list includes analyzing screenshots as part of a structured decision task, which is exactly this path.

Multimodal in practice, not just in tags: base64 screenshots in the state array, probabilities out.Watch at 3:25 - 9
The honest footer: 85.71% overall — and 49 GiB if you self-host
Back on the model card, the panel's Overall row: Jev 84.51%, Qwen3.8-27B 74.76%, pplx-decider-v1-27b 85.71%. That 85.71% is the number the narrator cites as beating Jev — a vendor-measured edge of 1.2 points on Perplexity's own 2,710-row panel. Credit where due: the card shown in the video keeps the disclaimer in plain sight — 'These are Perplexity-reported results; they do not establish that the model will outperform every real-world decision task.' Then the card states the other bill: local inference wants Python 3.12+ and a CUDA GPU with room for approximately 49 GiB of weights plus working memory. The $0.02/M hosted endpoint and the free-but-heavy weights are budgets for two different teams; the video's value is showing both doors in one sitting.

Vendor-reported 85.71% overall — with the vendor’s own disclaimer and the 49 GiB self-hosting bill in the same frame.Watch at 4:05 - 10
Six use cases Perplexity names
The card's Potential use cases list, exactly as shown: Customer support — automatically route tickets to the appropriate teams; AI agents — select actions or tools based on defined criteria; Business automation — classify requests and prioritize workflows; RAG evaluation — assess the reliability of retrieved answers; Content moderation — classify content against specified rules; Image-based decisions — analyze screenshots as part of a structured decision task. The narrator walks all six, and they share one shape: every row is a place where you would otherwise prompt an LLM for a label and parse a reply — the decision model returns the number directly, and your code thresholds it. The next two steps follow the list's first row into the official cookbook.

Six bullets, one shape: replace parse-the-reply with threshold-the-number.Watch at 4:11 - 11
The cookbook: triage support tickets, three questions each
The video's second half walks Perplexity's official recipe: Triage Support Tickets with the Decisions API — 'Route inbound support tickets with three probabilistic questions per ticket from the Decisions API, then send only the escalations to the Agent API.' The three questions, per the narrator: does this ticket need human intervention, what is the customer's intent, and how severe is the problem. Each ticket is one state; the three questions go out in one request; the returned probabilities and confidence then map each ticket to a route — escalate, review, queue, auto_reply or close, matching the escalated / needs-review / queued / auto-responded / closed outcomes the narrator lists. The docs page organizes the whole build stage by stage: the tickets, the script (three questions, one request; from probabilities to a route; one request per ticket), and a run section.

Three questions, one request per ticket — then only the escalations get expensive.Watch at 4:52 - 12
The run: 12 tickets in, 3 escalations out
The cookbook's Run it section, live: a twelve-row table with ticket, route, human_intent, conf, sev, P(crit) and reason columns. Three rows escalate: T-1001 (outage_or_bug, P(crit)=0.97, expected severity 2.97), T-1006 (security, 0.99, 2.99) and T-1007 (outage_or_bug, 0.85, 2.84) — the security breach and production failures the narrator points at. Everything else routes queue, auto_reply, close or review — including T-1011, sent to review not by severity but because its intent was unclear (confidence 0.55), which is the confidence number doing routing work. The summary line prints '3 of 12 escalated. Results: triage_results.jsonl. Add --draft-escalations to send the escalated tickets to the Agent API.' Timing, from the docs below the terminal: about 2.5 seconds for all twelve classification requests, most taking 150 to 250 ms at 590 to 659 input tokens per ticket — plan for hundreds of milliseconds per ticket, not tens. That is the whole economic argument in one screen: the cheap decision model reads everything, and the expensive Agent API only ever sees the three tickets that deserve it — the narrator's phrase is 'reducing unnecessary cloud calls to large models'.

12 tickets, about 2.5 seconds, 3 escalations: the cheap model reads everything, the Agent API sees three.Watch at 5:56
Frequently asked questions
What is PPLX Decider (pplx-decider-v1-27b)?
It is Perplexity's open-source decision model: an Apache 2.0 fine-tune of Qwen3.8-27B, published on Hugging Face at perplexity-ai/pplx-decider-v1-27b (Text Classification pipeline, multimodal, custom-code). Instead of generating text it returns probabilities — the probability of yes (noul), a distribution over your options (choice), or a level on your rubric (score). You can call it through the hosted Decisions API — which now defaults to the newer pplx-decider-v1.1-27b, with v1 still supported — or self-host the weights, which the model card sizes at roughly 49 GiB of GPU memory plus working memory on Python 3.12+ with CUDA.
How much does the Perplexity Decisions API cost?
The video quotes $0.04 per million input tokens with output tokens completely free. That input price has since been halved: as of October 7, 2026, the official pricing page lists $0.02 per million input tokens, output free, no per-request fee — the same figure our Jev-vs-Perplexity comparison recipe uses. At the cookbook's own measurement (590–659 input tokens per ticket), a twelve-ticket triage run costs a small fraction of a cent; the expensive calls are the escalations you forward to the Agent API, and the design's point is that there are only three of those out of twelve.
How does PPLX Decider compare to Jev?
Three frames from the video: on Perplexity's own 11-benchmark panel (September 2026) it edges Jev overall, 85.71% to 84.51%, winning FinancialPhraseBank and RAGTruth while Jev keeps WinoGrande, BBH and JudgeBench; on the community Jev Decision Index captured in the video, Jev leads 57.9 to 56.4 (rank #1 vs #4), though the open model's median latency is 101 ms against Jev's 524 ms hosted round trip; and structurally, Jev is closed while Decider ships Apache 2.0 weights. Note the board has since moved: Perplexity's v1.1 card reports 61.56 vs Jev's 57.9. Vendor-reported numbers, moving monthly — check the current cards before you quote any of them.
Can PPLX Decider take screenshots as input?
Yes — state can carry images next to text. Pass state as an array and put each image in an OpenAI-style image part with a base64 data URL; an image can also be the whole state. Constraints from the docs: PNG, JPEG and WebP data URLs only (the API never fetches http/https URLs — that returns 400), each image is read in 32×32-pixel tiles with a ceiling of 2,048 tiles (1440×1440 or 2048×1024 fit; 1600×1310 does not), oversized images stall about a minute and then return 504, and image tokens count as input at roughly 1,000 tokens per megapixel.
What are noul, choice and score in the Decisions API?
The three question types, all aimed at the same state in one request. noul: a yes/no check returning the probability of yes from 0 to 1. choice: pick one of your options — 1 to 255 — returning the most likely option, per-option probabilities and confidence. score: grade content on an ordered rubric of 2 to 10 levels, returning the probability-weighted level index, a legend, per-level probabilities and confidence. A request accepts 1 to 128 named questions under roughly 262,144 input tokens; confidence is the model's own certainty estimate rather than the top probability, and near-identical requests can differ in the second decimal, so threshold with margin.
Where does the 12-ticket triage example come from?
Perplexity's official cookbook, 'Triage Support Tickets with the Decisions API' — the page the video walks in its second half. Every ticket gets three probabilistic questions (needs human intervention? customer intent? severity?) in one request, and the probabilities plus confidence route it to escalate, review, queue, auto_reply or close. The documented run classifies twelve tickets in about 2.5 seconds (most requests 150–250 ms at 590–659 input tokens), escalates three (T-1001, T-1006, T-1007), and prints the flag --draft-escalations to forward exactly those to the Agent API — the 'only send escalations to the big model' economics the narrator highlights.
Related guides
Recipe: Jev vs Perplexity Decisions API
The compare-first step above this page: contracts, pricing and confidence quality side by side — read it to choose, come back here to use.
ReadDecision Model API Hub
The category hub: every hosted decision API and where the $0.02/M Decisions endpoint sits in the October 2026 field.
ReadNox-4B Decision Guide
Another open decision model, hands-on: the local-checkpoint route as an alternative to Perplexity’s hosted one.
ReadJev Pricing
What Jev charges per decision — the number the $0.02/M input rate has to beat for your workload.
ReadNanoJev Guide
Same release wave, opposite pole: a 0.6B replica you install and own, against this 28B model you call.
ReadJEV-27B Local Guide
A third-party bench test of another open 27B decision model — the testing methodology to reuse on Decider weights.
ReadMore video walkthroughs
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Jev Log Triage with Expanso Edge
- Jev Lead Enrichment with Treg: ICP and Signup Scoring
- Use Jev Decision Nodes in Heym for Model Routing
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)
- Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)
- Jev Context Compaction: Prune AI Agent Memory Without Generative Summaries
- Jev as an LLM Judge: Confidence-Gated Cascades at 0.36% of the Cost
- TypeSafe Computer Use: Local Desktop Automation with Jev, Step by Step
- Jev Resume Screening: Build an AI Resume Evaluator with the Jev JavaScript SDK
- Jev + Claude Code Guide: Voice-Controlled Browser Automation with Typed Decisions
- Jev + Codex: Install the TypeSafe Skill and Triage a Real Gmail Inbox
- Jev vs Ollama: Can Local AI Replace Hosted Jev Without Sending Your Data Away?
- Build Your Own Jev: Train a Free Open-Source Zero-Shot Classifier (That Plays Doom)
- CUA-S1-Forms: a 706K-Parameter Jev-Like Model That Fills GUI Forms on Your CPU
- NOC/SOC Alert Triage with Jev: Rules First, One Typed Question, a Policy Gate
- Ollama Decision Models: Run tev1 and Nimble Locally (Tested on an 8 GB Card)
- Jev Guardrails in Production: A Five-Step Playbook for Decision Automation
- A Session Drift Guard for Pi Agent: Let the Jev Model Propose, Let Code Decide
- 50 Tev1 Use Cases: What a Local Decision Model Can Actually Do (Tested on an 8 GB Laptop)
- Jev, Hands-On: Where the Official Claims Meet Independent Remeasurement
- Jev Ticket Classification in a Real App: the After-Insert Hook and the Calculated Field
- OpenJev RLCD: Run the Open-Source Calibrated Decision Model Locally (Full Guide)
- Clef-Flash vs Jev: I Tested Cloudflare's Decision Model Locally in Ollama (Q4_K_M)
- Clef-Flash Tutorial: Install and Run Cloudflare's 9B Multimodal Decision Model Locally on Ubuntu
- CLM-8B: the Contrastive Decision Model That Scores 1,024 Options in 44 ms (13x Faster Than Jev)
- Nox 4B Tutorial: Run the Decision 2.0 Model Locally and Put It Through Four Real Decisions (One Ends in a Fail)
- Julia-1 Tutorial: Install the Open-Source Jev Replacement in Pure Python (and Watch It Beat If-Statements 9 to 2)
- OpenJev 0.8B on CPU: I Built a Ticket-Routing Inbox and Calibrated the Thresholds
- Jev n8n Integration: the JevGate Community Node, Step by Step (Plus a Plain-HTTP Fallback)
- NanoJev Tutorial: Install the 0.6B Open-Source Jev Replica and Watch It Route Decisions at 47 ms
- AutoTrust JEV-27B Tested Locally: Four Arms, 72 Cases, and One 0.969 Score That Was Wrong