Jev recipe / vendor comparison

Jev vs Perplexity Decisions API: GA, Priced, and Compared

Perplexity answered Jev with a GA decision API and published pricing. Compare the verified contracts: the same noul/choice/score primitives, $0.04 vs $0.042 per million input tokens, Apache 2.0 weights, vision input — and the details that still decide the switch.

AD

Six checks before you pick a decision API

01Read this as the second entrant in one category, not a new debate. Perplexity’s Decisions API is the first decision-model API to ship GA and published pricing together — OpenAI’s Luna-based Decisions API is still preview with no separately priced line. The three primitives are named exactly like Jev’s: noul, choice, score. Any Perplexity API key works. When the surface vocabulary converges this far, the comparison moves off naming and onto contract details, pricing, deployment, and calibration. Everything below was verified against official documentation on 2026-10-06, one day into launch week.
02Map the output contracts cell by cell — same names, different details. Both APIs return typed answers keyed by question name, with no free-text generation. Jev’s decision heads return an answer plus an RLCD-calibrated confidence. Perplexity returns a noul probability for yes/no, per-option probabilities plus a most-likely choice plus a confidence for choice, and — for score — a probability-weighted average of rubric level indices with a legend. Two documented details matter if you automate: Perplexity’s confidence is “the model’s own certainty estimate,” not the top probability, and identical requests usually return identical numbers but can differ in the second decimal place. If a 0.01 wobble or a non-top-probability confidence changes your gate behavior, that is the comparison.
03Price the triangle honestly — the old differentiator is gone. Jev publishes $0.042 per million input tokens with output free. Perplexity publishes $0.04 per million input tokens, output free, no per-request fee — 5% cheaper on input, the same “output does not exist” philosophy. The third point, OpenAI’s gpt-6-luna, bills $0.10 per million input and $0.50 per million output at short context. So input-only billing no longer separates Jev from anyone: Perplexity matched it on day one. What remains is rate shape (Perplexity documents 10 requests/second on every plan), image tokens (about 1,000 per megapixel), and the axes pricing cannot reach — independent benchmarks, calibration methodology, and tooling.
04Check the modality and capacity edges — Perplexity has one capability Jev does not. Vision input: the Decisions API accepts images in state as OpenAI-style image parts (base64 PNG/JPEG/WebP data URLs only; the API never fetches a URL), read in 32×32 pixel tiles with up to 2,048 tiles per image — 1440×1440 or 2048×1024 fit — and an oversized image does not fail fast: it waits about a minute and returns 504. Capacity: 262,144 input tokens versus Jev’s 32k context, up to 128 questions per request, 255 options per choice, 10 score levels. If your decisions read screenshots or documents today, that row wins by itself — we state the same plain fact on the Clef page. If everything you decide is text, Jev’s SDK, playground, and multi-question evaluate cover the same ground.
05Treat “same primitives” as an A/B hypothesis, not a port. Same names are not the same contract: Perplexity requires model: "pplx-decider-v1-27b" on every request (missing or unknown returns 400), nests answers under answers, authenticates with a Bearer header only (an x-api-key header is not read), rejects unknown top-level fields, and serves decisions at POST /v1/decisions with no trailing slash. Wrap both endpoints behind one decision interface, then re-validate the swap on ~100 of your own labeled examples before any auto-action — especially across the confidence semantics from step two. Design for the documented rate limit too: 10 requests/second on every plan, with 429s carrying Retry-After.
06Plan the mixed deployment and the revisit cadence from day one. Keep both vendors behind one gate, let the confidence-gated fallback chain route low-confidence traffic from either side into the same review queue, and re-check this page monthly: the API went GA in launch week (announced 2026-10-05; this page verified the docs on 2026-10-06), no independent benchmark of pplx-decider exists yet, and launch-week numbers move. Every latency and quality figure above is vendor-documented, labeled as such, and should be re-tested on your own workload before it drives an architecture.
schema / typed multi-question contract
{
  "routing": {
    "type": "choice",
    "instructions": "Which queue should this ticket go to?",
    "criteria": {
      "billing": "Payment, invoice or refund issue",
      "technical": "Product malfunction or bug",
      "sales": "Buying or upgrade question",
      "abuse": "Safety or abuse report"
    }
  },
  "needs_safety_escalation": {
    "type": "noul",
    "instructions": "Does this ticket require a safety escalation regardless of queue?"
  },
  "urgency": {
    "type": "score",
    "instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)"
  }
}
This schema is the contract depth in one view — run all three questions in the Playground and inspect the calibrated confidence.

Decision API scorecards: TypeSafe Jev vs Perplexity Decisions API (GA launch week, sources labeled per cell)

METRIC
Jev
Perplexity Decisions API
Output contract
Native decision heads

a typed answer (Choice/Score/Noul) with an RLCD-calibrated probability; nothing generated; deterministic single-pass semantics

Typed answers under answers, keyed by question name: noul probability, choice with per-option probabilities + confidence, score as a probability-weighted average of level indices with a legend

the docs note identical requests “occasionally differ in the second decimal place”

Latency & throughput (labeled sources)
~70–100ms typical single-pass evaluate call

our own 49-task P95 series is public at /benchmarks; text state, 32k context

Vendor-tested response time grows with input: under 2s for small requests, ~5s at ~90k tokens, ~14s at ~190k, ~23s near the 262k limit (Perplexity tests, Sep 30, 2026); 10 requests/second on every plan; no independent benchmark exists yet
Pricing (October 2026, the triangle)
Public: $0.042 per million input tokens, output free

pennies per 100k decisions

Public: $0.04 per million input tokens, output free, no per-request fee

5% under Jev on input; the third point is OpenAI gpt-6-luna at $0.10 in / $0.50 out (short context), Decisions API not priced separately

Deployment & license
Hosted TypeSafe API + SDK

proprietary hosted contract; the local route runs through third-party OpenJev clones

Hosted GA API + Apache 2.0 open weights on Hugging Face (perplexity-ai/pplx-decider-v1-27b, Qwen3.8-27B base) with official Python inference code

a vendor-sanctioned self-host path; you provision the serving stack and own its validation

Calibration story
RLCD-trained probabilities

a 0.85 threshold behaves like an accuracy contract; ECE methodology published on our benchmarks page

Confidence documented as “the model’s own certainty estimate”

explicitly not the top probability, and it drops when the runner-up option closes in; no published calibration/ECE methodology yet, so thresholds need your own labels from day one

Ecosystem & modality
Text-only state, 32k context, multiple questions per call, SDK + Playground, plus the audit-trail / fallback / calibration content stack around it
Vision input today (OpenAI-style image parts, 32×32 tiles, ≤2,048 tiles per image, ~1,000 tokens per megapixel), 262k input tokens, 128 questions per request, 255 options, 10 score levels, 32 MiB bodies; sits next to Perplexity’s Agent and Router APIs

Side-by-side code: the Jev typed call vs the Perplexity decisions endpoint and the A/B switch

python / jev typed multi-question evaluate
import requests

JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_ROUTE_CONFIDENCE = 0.85

QUESTIONS = {
    "routing": {
        "type": "choice",
        "instructions": "Which queue should this ticket go to?",
        "criteria": {
            "billing": "Payment, invoice or refund issue",
            "technical": "Product malfunction or bug",
            "sales": "Buying or upgrade question",
            "abuse": "Safety or abuse report",
        },
    },
    "needs_safety_escalation": {
        "type": "noul",
        "instructions": "Does this ticket require a safety escalation regardless of queue?",
    },
    "urgency": {
        "type": "score",
        "instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
    },
}

resp = requests.post(
    JEV_ENDPOINT,
    json={"state": {"ticket": "..."}, "questions": QUESTIONS},
    timeout=5,
)
resp.raise_for_status()
data = resp.json()

route = data["routing"]  # typed answer + calibrated confidence, no generation
if (
    route["confidence"] >= AUTO_ROUTE_CONFIDENCE
    and not data["needs_safety_escalation"]["answer"]
):
    lane = f"auto:{route['answer']}"  # single pass, input tokens only
else:
    lane = "review"  # 0.60-0.85 band or safety flag
python / the Perplexity decisions endpoint (GA contract, as documented)
import os

import requests

# POST /v1/decisions - no trailing slash (a trailing slash returns 404).
# Auth is Bearer only: a key in x-api-key is not read and returns 401.
PPLX_ENDPOINT = "https://api.perplexity.ai/v1/decisions"
AUTO_ROUTE_CONFIDENCE = 0.85

resp = requests.post(
    PPLX_ENDPOINT,
    headers={"Authorization": f"Bearer {os.environ['PERPLEXITY_API_KEY']}"},
    json={
        # model is REQUIRED on every request; missing/unknown returns 400
        "model": "pplx-decider-v1-27b",
        "state": {
            "ticket": "Checkout has been failing for every customer "
            "for the last hour."
        },
        "questions": {
            "routing": {
                "type": "choice",
                "instructions": "Which queue should this ticket go to?",
                "criteria": {
                    "billing": "Payment, invoice or refund issue",
                    "technical": "Product malfunction or bug",
                    "sales": "Buying or upgrade question",
                    "abuse": "Safety or abuse report",
                },
            },
            "needs_safety_escalation": {
                "type": "noul",
                "instructions": "Does this ticket require a safety "
                "escalation regardless of queue?",
            },
            "urgency": {
                "type": "score",
                "instructions": "Rate how urgent a human reply is, "
                "routine to business-stopping",
                # score: 1-10 ordered levels, answered with the
                # probability-weighted average of level indices (0-based)
                "criteria": ["Routine", "Minor", "Elevated", "Urgent",
                             "Business-stopping"],
            },
        },
    },
    timeout=30,  # docs: small inputs answer under 2s; 30s covers the limit
)
resp.raise_for_status()
data = resp.json()

answers = data["answers"]  # keyed by question name, typed like its question
route = answers["routing"]  # choice + confidence + per-option probabilities
if (
    route["confidence"] >= AUTO_ROUTE_CONFIDENCE
    and answers["needs_safety_escalation"]["noul"] < 0.5  # P(yes)
):
    lane = f"auto:{route['choice']}"
else:
    lane = "review"  # low confidence or safety flag

# Honest-boundary notes from the docs:
# - confidence is "the model's own certainty estimate", NOT the top
#   probability; it drops when the runner-up option closes in.
# - identical requests usually return identical numbers, but they can
#   differ in the second decimal place - validate before automating.
# - usage.input_tokens is billed at $0.04/M; output_tokens are free.
# - rate limit: 10 requests/second on every plan (429 + Retry-After).
typescript / jev typed multi-question call
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_ROUTE_CONFIDENCE = 0.85;

const QUESTIONS = {
  routing: {
    type: "choice",
    instructions: "Which queue should this ticket go to?",
    criteria: {
      billing: "Payment, invoice or refund issue",
      technical: "Product malfunction or bug",
      sales: "Buying or upgrade question",
      abuse: "Safety or abuse report",
    },
  },
  needs_safety_escalation: {
    type: "noul",
    instructions:
      "Does this ticket require a safety escalation regardless of queue?",
  },
  urgency: {
    type: "score",
    instructions:
      "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
  },
} as const;

const resp = await fetch(JEV_ENDPOINT, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ state: { ticket: "..." }, questions: QUESTIONS }),
});
const data = await resp.json();

const lane =
  data.routing.confidence >= AUTO_ROUTE_CONFIDENCE &&
  !data.needs_safety_escalation.answer
    ? `auto:${data.routing.answer}`
    : "review";
typescript / one gate, two backends (the shape-aware A/B switch)
// Same primitive names, different contracts - adapt both shapes behind
// one interface. Differences this adapter absorbs: Perplexity requires
// model: "pplx-decider-v1-27b" on every call (400 otherwise), nests
// answers under `answers`, uses Bearer auth only, rejects unknown fields
// and trailing slashes, and documents 10 requests/second on every plan.
// Its confidence is the model's own certainty estimate, not the top
// probability - validate calibration per vendor before automating on it.
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const PPLX_ENDPOINT = "https://api.perplexity.ai/v1/decisions";

const QUESTIONS = {
  routing: {
    type: "choice",
    instructions: "Which queue should this ticket go to?",
    criteria: {
      billing: "Payment, invoice or refund issue",
      technical: "Product malfunction or bug",
      sales: "Buying or upgrade question",
      abuse: "Safety or abuse report",
    },
  },
} as const;

type TypedAnswer = { answer: string; confidence: number };

export async function routeTicket(ticket: string): Promise<TypedAnswer> {
  if (process.env.DECISION_BACKEND === "perplexity") {
    const resp = await fetch(PPLX_ENDPOINT, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${process.env.PERPLEXITY_API_KEY!}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify({
        model: "pplx-decider-v1-27b", // required on every request
        state: { ticket },
        questions: QUESTIONS,
      }),
    });
    if (!resp.ok) throw new Error("decision call failed");
    const data = await resp.json();
    const a = data.answers.routing as {
      choice: string;
      confidence: number;
      probabilities: Record<string, number>;
    };
    return { answer: a.choice, confidence: a.confidence };
  }

  const resp = await fetch(JEV_ENDPOINT, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ state: { ticket }, questions: QUESTIONS }),
  });
  if (!resp.ok) throw new Error("decision call failed");
  const data = await resp.json();
  return data.routing as TypedAnswer;
}

Jev vs Perplexity Decisions API FAQ

What is the Perplexity Decisions API?

A GA decision-model API from Perplexity, announced during launch week (2026-10-05) and verified in the official docs on 2026-10-06: send a state (a string, object, or array — images included) plus 1–128 named questions, and get probabilities back instead of generated text. One model serves it, pplx-decider-v1-27b, and the three question types are named exactly like Jev’s primitives: noul (probability of yes), choice (per-option probabilities plus confidence), and score (a probability-weighted average on a rubric of up to 10 levels). Pricing is $0.04 per million input tokens with output tokens free and no per-request fee, and the weights are Apache 2.0 on Hugging Face.

How is it different from Jev?

Same category, same primitive names, different contract details. Jev returns a typed answer with an RLCD-calibrated confidence and deterministic single-pass behavior; Perplexity documents confidence as “the model’s own certainty estimate” (not the top probability) and says identical requests can differ in the second decimal place. Perplexity accepts image input and up to 262k input tokens where Jev is text-only with a 32k context; Jev publishes calibration methodology (ECE on our benchmarks page) and ships an SDK and playground. Both bill input-only with output free — $0.042/M for Jev, $0.04/M for Perplexity.

Is there an open source alternative to the Perplexity Decisions API?

Perplexity itself is one: the weights (perplexity-ai/pplx-decider-v1-27b on Hugging Face) are Apache 2.0 with official Python inference code, so you can run predictions on infrastructure you operate — you provision the serving stack and own the validation. The Jev ecosystem’s local route runs through the OpenJev-family clones (Kev, SemIf, Von), and Cloudflare’s Clef is the other open-weights decision-model family. The honest caveat: self-hosting reproduces the weights, not necessarily the hosted API’s full behavior — image tile handling, confidence semantics, and calibration on your stack are yours to verify.

Which one should I pick?

By scenario. Perplexity-first: your decisions read images today, you need more than 32k context, you want open weights with a vendor-sanctioned self-host path, or extreme-volume input cost is the deciding line ($0.04 vs $0.042). Jev-first: you automate on confidence thresholds and need published calibration methodology, deterministic behavior for replayable audits, the SDK + playground workflow, or the audit-trail and fallback content stack. Either way, the deciding data is local: run ~100 of your own labeled examples through both gates — launch-week vendor docs answer a scouting question, not a deployment one.

Is the $0.04 vs $0.042 difference the real cost story?

No — at sane volumes both are pennies. Six hundred input tokens cost about $0.000024 on Perplexity and $0.000025 on Jev; the 5% gap only matters at nine-figure monthly token counts. The cost axes that actually move money: images (about 1,000 input tokens per megapixel on Perplexity), output billing (free on both — that is what separates both from OpenAI’s gpt-6-luna at $0.10 in / $0.50 out), and rate limits (Perplexity documents 10 requests/second on every plan). Model the triangle with the cost calculator, then stop optimizing pennies and start validating confidence.

Can the two APIs share one request path?

Behind an adapter, yes; as a drop-in, no. The primitive names and the state-plus-questions philosophy match, but Perplexity requires model: "pplx-decider-v1-27b" on every call, nests responses under answers, authenticates with a Bearer header only (x-api-key is not read), rejects unknown fields and trailing slashes, and documents 10 requests per second on every plan. Put both behind one decision interface like the two-backend code above, keep the confidence semantics per vendor, and validate the swap on labeled examples before automating.

Go deeper on the decision-API cluster