Jev recipe / document processing

Document Classification with Jev: A Two-Stage Packet Pipeline

One email with three attachments, one scan mixing eight invoice pages with a contract: split the packet with Noul boundary questions, then classify each document with a typed Choice — calibrated confidence routes the review queue.

Build the packet pipeline in six steps

01Define both axes before writing code. The packet axis: what starts a new document — a new invoice number, a signature block, a change in layout. The type axis: invoice, contract, receipt, letter, other. Write both as criteria vocabulary your ops team already uses when they file documents by hand.
02Keep the state small and textual: the page text from your extractor plus decision-relevant metadata — page_index, source (email, scanner, fax), whether OCR ran. Jev reads text, not pixels: there is no OCR or layout parser inside the model, so extraction happens upstream and its output is the input.
03Run the two stages as one evaluate call per page: Noul questions decide packet boundaries (is_single_document, needs_ocr), and the doc_type Choice labels the segment that page belongs to. Every answer arrives with calibrated confidence in a single ~100–500ms forward pass — pages run in parallel, so wall time stays flat as packets grow.
04Own the routing policy in code: doc_type confidence ≥ 0.85 auto-files, 0.60–0.85 drops into a review queue, and below 0.60 — plus every "other" label and every needs_ocr page — escalates to a human. The same three-lane discipline as the confidence-gated fallback chain: cutoffs derived from labeled packets, never guessed.
05Do the batch math before committing: states bill at $0.042/M input tokens and output is free. A 9-page packet at ~300 tokens per page costs ≈ $0.0001, so a 10,000-packet back-catalog scan lands near a dollar. The ecosystem's published datapoint: 1,018 research papers classified for $0.08 total.
06Monitor boundaries, not just labels: track how often is_single_document lands in the contested band, route ops corrections into a labeled set, and re-validate thresholds quarterly. A new vendor's document layout is a distribution shift — the same drift the fallback chain recalibrates for.
schema / packet split & classify contract
{
  "is_single_document": {
    "type": "noul",
    "instructions": "Does this page start a new document, or continue the document on the previous page?"
  },
  "needs_ocr": {
    "type": "noul",
    "instructions": "Is this page's text missing or too garbled to decide on, such that OCR must run before a reliable call?"
  },
  "doc_type": {
    "type": "choice",
    "instructions": "Classify this document into exactly one type",
    "criteria": {
      "invoice": "Itemized billing: line items, quantities, totals, tax, payment terms",
      "contract": "Binding agreement: parties, clauses, term dates, signature blocks",
      "receipt": "Proof of completed payment: paid marker, transaction reference, nothing due",
      "letter": "Human correspondence: salutation, body prose, signature block",
      "other": "None of the above matches with confidence — escalate instead of forcing a label"
    }
  }
}
Send a real page as state and inspect is_single_document, needs_ocr, and doc_type with calibrated confidence.

The document type axis: five labels and a designed fallback

Classification quality starts with the label set. The five types below cover most back-office volume; the sixth card is the most important one — a designed "other" that escalates instead of forcing a bad label. Use the tells as criteria vocabulary in your Choice schema, and keep the axis small: every extra label dilutes the probability mass.

Invoices

Money owed — the packet's most common payload.

  • Line items with quantities and unit prices
  • An invoice number in a header or footer
  • Tax lines and a total-due amount
  • Payment terms: net 30, due date, remit-to
  • Issuer and bill-to blocks

Contracts

Binding agreement — high stakes, worth a review lane.

  • Named parties in a preamble
  • Numbered clauses and a definitions section
  • Term, effective, and renewal dates
  • Signature blocks with printed names and dates
  • Boilerplate legal headers and page numbering

Receipts

Payment proof — most often confused with invoices.

  • A paid or transaction-completed marker
  • Terminal or POS header with store details
  • Short line items with no payment terms
  • Last-four card digits or a payment reference
  • Nothing left due — amount paid, zero balance

Letters & correspondence

Human communication — cover emails, memos, notices.

  • Salutation and sign-off around free-form prose
  • A request or announcement in body text
  • Signature block with title and contact details
  • Quoted reply chains in email bodies
  • No billing fields anywhere

Forms & intake

Structured capture — applications, claims, onboarding.

  • Labeled fields with values in fixed positions
  • Checkboxes, dropdowns, and option markers
  • Section headers numbered by the form design
  • Handwriting or signature boxes
  • Page N of M footers

Other — the designed fallback

A label, not a failure: route it, don't force it.

  • Documents that match no criterion above
  • Mixed or empty extraction output
  • Confidence below threshold on every label
  • Unknown layouts from a new vendor
  • Treated as a work queue, not an error
Interactive demo / packet split & classify

Two-stage packet split & classification simulator

Pick a packet and run the two-stage pipeline: per-page Noul boundaries, then segments, then Choice types with confidence lanes (front-end simulation only, no API calls):

is_single_document + needs_ocr + doc_type
4 pages
p1

From: ops@acme.co — "Hi team, attached are this month's billing docs: the March invoice, the card receipt, and the signed MSA renewal."

p2

INVOICE #2026-0341 · ACME Supplies Ltd. · 14 line items · Subtotal $11,618.32 · VAT 7.4% · Total due: $12,480.00 · Net 30

p3

RECEIPT — PAID · Visa ****4218 · $12,480.00 · Ref TXN-88213 · 2026-03-02 · Nothing due

p4

MASTER SERVICES AGREEMENT · Acme Co. × Beta LLC · Term 2026-04-01 – 2028-03-31 · Signatures: [executed]

source: inbound / ops@acme.co
Jev decision output
~148ms / ≈$0.0001

Press "Run split + classify" to see boundary verdicts, packet segments, and per-document type confidence.

Policy in code: is_single_document = true -> new segment · needs_ocr || doc_type = other -> human · confidence >= 0.85 -> auto-file · 0.60–0.85 -> review queue · < 0.60 -> human review

Packet classification: Jev two-stage pipeline vs one long LLM prompt

METRIC
Jev
One long generative LLM prompt
Packet splitting accuracy
Per-page Noul with calibrated confidence

contested boundaries escalate to review instead of guessing

One long-prompt guess for the whole packet

wrong splits arrive looking confident

Type label reliability
Typed Choice over versioned criteria

0% schema errors, nothing to parse

Free-text verdict

2–8% malformed JSON and parse-and-retry loops

Latency per packet
~100–500ms per page, pages in parallel

a 9-page scan stays under ~500ms wall time

2–6s generating verdicts token-by-token across every page
Batch cost
≈$0.0001 per 9-page packet

input-only billing at $0.042/M, output free

Input bill again per page plus output tokens for every verdict and retry
Hallucination risk
Zero generation

the model cannot invent page text, amounts, or labels outside the schema

Can paraphrase or fabricate line items; every extracted amount needs separate validation
Mixed-layout robustness
Per-page independence

one garbled page degrades one decision, flagged by needs_ocr

Context dilution

nine pages of noise bury the one clause that matters

Production code: the packet pipeline in TypeScript and Python

typescript / two-stage packet split and classify
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";

// Policy lives in application code — the model only reports facts.
const AUTO_FILE_CONFIDENCE = 0.85; // >= 0.85 -> file automatically
const REVIEW_MIN_CONFIDENCE = 0.6; // 0.60-0.85 -> review, below -> human

const QUESTIONS = {
  is_single_document: {
    type: "noul",
    instructions:
      "Does this page start a new document, or continue the document on the previous page?",
  },
  needs_ocr: {
    type: "noul",
    instructions:
      "Is this page's text missing or too garbled to decide on, such that OCR must run first?",
  },
  doc_type: {
    type: "choice",
    instructions: "Classify this document into exactly one type",
    criteria: {
      invoice: "Itemized billing: line items, quantities, totals, tax, payment terms",
      contract: "Binding agreement: parties, clauses, term dates, signature blocks",
      receipt: "Proof of completed payment: paid marker, transaction reference, nothing due",
      letter: "Human correspondence: salutation, body prose, signature block",
      other: "None of the above matches with confidence",
    },
  },
} as const;

export type DocLane = "auto_file" | "review" | "human_review";

export async function evaluatePage(page: {
  text: string;
  page_index: number;
  source: string;
  has_ocr: boolean;
}) {
  const response = await fetch(JEV_ENDPOINT, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ state: page, questions: QUESTIONS }),
  });
  if (!response.ok) throw new Error("Jev evaluate failed");
  const data = await response.json();

  const { is_single_document, needs_ocr, doc_type } = data;
  const lane: DocLane =
    needs_ocr.answer || doc_type.answer === "other"
      ? "human_review"
      : doc_type.confidence >= AUTO_FILE_CONFIDENCE
        ? "auto_file"
        : doc_type.confidence >= REVIEW_MIN_CONFIDENCE
          ? "review"
          : "human_review";

  return {
    page_index: page.page_index,
    starts_document: is_single_document.answer,
    boundary_confidence: is_single_document.confidence,
    needs_ocr: needs_ocr.answer,
    doc_type: doc_type.answer,
    doc_type_confidence: doc_type.confidence,
    lane,
  };
}

// Boundary and type answers come back in the same call: pages where
// starts_document = true open a new segment; the rest accumulate.
export function splitPacket(pages: Awaited<ReturnType<typeof evaluatePage>>[]) {
  const segments: (typeof pages)[] = [];
  for (const page of pages) {
    if (page.starts_document || segments.length === 0) segments.push([page]);
    else segments[segments.length - 1].push(page);
  }
  return segments;
}

// Pages are independent — fan them out concurrently. A 9-page packet
// stays under ~500ms wall time, input tokens only, 0% schema errors.
const results = await Promise.all(packetPages.map(evaluatePage));
const segments = splitPacket(results);
const reviewQueue = results.filter((r) => r.lane !== "auto_file");
python / batch classification for an inbound scan queue
import requests

JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_FILE_CONFIDENCE = 0.85   # >= 0.85 -> file automatically
REVIEW_MIN_CONFIDENCE = 0.60  # 0.60-0.85 -> review queue, below -> human

QUESTIONS = {
    "is_single_document": {
        "type": "noul",
        "instructions": "Does this page start a new document, or continue the document on the previous page?",
    },
    "needs_ocr": {
        "type": "noul",
        "instructions": "Is this page's text missing or too garbled to decide on, such that OCR must run first?",
    },
    "doc_type": {
        "type": "choice",
        "instructions": "Classify this document into exactly one type",
        "criteria": {
            "invoice": "Itemized billing: line items, quantities, totals, tax, payment terms",
            "contract": "Binding agreement: parties, clauses, term dates, signature blocks",
            "receipt": "Proof of completed payment: paid marker, transaction reference, nothing due",
            "letter": "Human correspondence: salutation, body prose, signature block",
            "other": "None of the above matches with confidence",
        },
    },
}

def evaluate_page(page: dict) -> dict:
    resp = requests.post(
        JEV_ENDPOINT,
        json={"state": page, "questions": QUESTIONS},
        timeout=5,
    )
    resp.raise_for_status()
    data = resp.json()
    doc = data["doc_type"]

    if data["needs_ocr"]["answer"] or doc["answer"] == "other":
        lane = "human_review"
    elif doc["confidence"] >= AUTO_FILE_CONFIDENCE:
        lane = "auto_file"
    elif doc["confidence"] >= REVIEW_MIN_CONFIDENCE:
        lane = "review"
    else:
        lane = "human_review"

    return {
        "page_index": page["page_index"],
        "starts_document": data["is_single_document"]["answer"],
        "doc_type": doc["answer"],
        "confidence": doc["confidence"],
        "lane": lane,
    }

def classify_packet(pages: list[dict]) -> list[list[dict]]:
    # Pages are independent — fan out concurrently in production.
    results = [evaluate_page(p) for p in pages]
    segments: list[list[dict]] = []
    for page in results:
        if page["starts_document"] or not segments:
            segments.append([page])
        else:
            segments[-1].append(page)
    return segments

def scan_inbound_queue(packets: list[list[dict]]) -> dict:
    # ~100-500ms per page, input-only pricing at $0.042/M: a
    # 10,000-packet back-catalog scan costs about a dollar, no GPUs.
    lanes: dict[str, int] = {}
    for packet in packets:
        for segment in classify_packet(packet):
            lane = segment[0]["lane"]
            lanes[lane] = lanes.get(lane, 0) + 1
    return lanes

Document classification FAQ

How is document classification different from text classification?

Text classification labels one self-contained text — one ticket, one review, one message. Document classification faces packets: an email with attachments, a scanner batch that mixed eight invoice pages with a contract, a PDF export that concatenated unrelated documents. The two-stage pipeline adds the step text classification never needs — deciding where one document ends and the next begins with a per-page Noul — and only then classifies each segment with a typed Choice. The single-text general case is covered in the text classification API guide.

Can Jev read scanned PDFs, images, or handwriting?

No — and being upfront about this is the difference between a demo and a pipeline. Jev's state is text: there is no OCR, no layout analysis, no vision inside the model. Scanned packets run OCR and extraction upstream and pass the resulting text in. What Jev adds is a needs_ocr Noul that flags pages whose text is missing or too garbled to decide on, so they land in a repair queue instead of getting a confident wrong label. Garbled OCR caps the accuracy of every downstream classifier — typed or generative.

How does the packet splitting actually work?

Each page gets an is_single_document Noul: "does this page start a new document, or continue the previous one?" Pages where the answer is true (always the first page) open a segment; the other pages accumulate into the open segment. The calibrated confidence matters more than the boolean: a boundary at 0.55 confidence is a contested boundary, and contested boundaries — like every other low-confidence call — can drop into the review queue instead of silently separating a contract from its invoice.

What confidence thresholds should I use?

Start with the same three lanes as the confidence-gated fallback chain: ≥ 0.85 auto-files, 0.60–0.85 goes to a review queue, and below 0.60 — plus every "other" label and every needs_ocr page — escalates to a human. Then derive your own cutoffs from labeled samples: plot predicted confidence against ops verdicts on real packets and place each cutoff where a mis-filed document is cheap to catch. Shadow-mode the policy before enforcement and re-validate quarterly — new vendors mean new layouts mean distribution shift.

What is DocJev — and is it official?

DocJev is an open-source library by Jerry Liu, the founder of LlamaIndex, that classifies and splits complex document packets — the same two-stage pattern this page implements. It is featured on the madewithjev showcase under Documents & OCR, with the author reporting ~139ms per packet and 40/40 accuracy on his own benchmark (author-reported, not an independent audit). DocJev is a third-party ecosystem project, not an official TypeSafe AI product — and neither is madewithjev.

What does classifying a large batch cost?

Billing is input tokens only at $0.042/M and output is free. A 9-page packet at roughly 300 tokens per page is ~2,700 input tokens ≈ $0.0001, so a 10,000-packet back-catalog scan lands near a dollar with no GPUs. A published ecosystem datapoint: 1kpapers classified 1,018 AI research papers for $0.08 total at ~256ms median latency. Generative pipelines pay a comparable input bill plus output tokens for every verdict — and again for every retry.

Extend the document pipeline