Jev recipe / document processing
Document Classification with Jev: A Two-Stage Packet Pipeline
One email with three attachments, one scan mixing eight invoice pages with a contract: split the packet with Noul boundary questions, then classify each document with a typed Choice — calibrated confidence routes the review queue.
Build the packet pipeline in six steps
{
"is_single_document": {
"type": "noul",
"instructions": "Does this page start a new document, or continue the document on the previous page?"
},
"needs_ocr": {
"type": "noul",
"instructions": "Is this page's text missing or too garbled to decide on, such that OCR must run before a reliable call?"
},
"doc_type": {
"type": "choice",
"instructions": "Classify this document into exactly one type",
"criteria": {
"invoice": "Itemized billing: line items, quantities, totals, tax, payment terms",
"contract": "Binding agreement: parties, clauses, term dates, signature blocks",
"receipt": "Proof of completed payment: paid marker, transaction reference, nothing due",
"letter": "Human correspondence: salutation, body prose, signature block",
"other": "None of the above matches with confidence — escalate instead of forcing a label"
}
}
}The document type axis: five labels and a designed fallback
Classification quality starts with the label set. The five types below cover most back-office volume; the sixth card is the most important one — a designed "other" that escalates instead of forcing a bad label. Use the tells as criteria vocabulary in your Choice schema, and keep the axis small: every extra label dilutes the probability mass.
Invoices
Money owed — the packet's most common payload.
- Line items with quantities and unit prices
- An invoice number in a header or footer
- Tax lines and a total-due amount
- Payment terms: net 30, due date, remit-to
- Issuer and bill-to blocks
Contracts
Binding agreement — high stakes, worth a review lane.
- Named parties in a preamble
- Numbered clauses and a definitions section
- Term, effective, and renewal dates
- Signature blocks with printed names and dates
- Boilerplate legal headers and page numbering
Receipts
Payment proof — most often confused with invoices.
- A paid or transaction-completed marker
- Terminal or POS header with store details
- Short line items with no payment terms
- Last-four card digits or a payment reference
- Nothing left due — amount paid, zero balance
Letters & correspondence
Human communication — cover emails, memos, notices.
- Salutation and sign-off around free-form prose
- A request or announcement in body text
- Signature block with title and contact details
- Quoted reply chains in email bodies
- No billing fields anywhere
Forms & intake
Structured capture — applications, claims, onboarding.
- Labeled fields with values in fixed positions
- Checkboxes, dropdowns, and option markers
- Section headers numbered by the form design
- Handwriting or signature boxes
- Page N of M footers
Other — the designed fallback
A label, not a failure: route it, don't force it.
- Documents that match no criterion above
- Mixed or empty extraction output
- Confidence below threshold on every label
- Unknown layouts from a new vendor
- Treated as a work queue, not an error
Two-stage packet split & classification simulator
Pick a packet and run the two-stage pipeline: per-page Noul boundaries, then segments, then Choice types with confidence lanes (front-end simulation only, no API calls):
From: ops@acme.co — "Hi team, attached are this month's billing docs: the March invoice, the card receipt, and the signed MSA renewal."
INVOICE #2026-0341 · ACME Supplies Ltd. · 14 line items · Subtotal $11,618.32 · VAT 7.4% · Total due: $12,480.00 · Net 30
RECEIPT — PAID · Visa ****4218 · $12,480.00 · Ref TXN-88213 · 2026-03-02 · Nothing due
MASTER SERVICES AGREEMENT · Acme Co. × Beta LLC · Term 2026-04-01 – 2028-03-31 · Signatures: [executed]
Press "Run split + classify" to see boundary verdicts, packet segments, and per-document type confidence.
Packet classification: Jev two-stage pipeline vs one long LLM prompt
contested boundaries escalate to review instead of guessing
wrong splits arrive looking confident
0% schema errors, nothing to parse
2–8% malformed JSON and parse-and-retry loops
a 9-page scan stays under ~500ms wall time
input-only billing at $0.042/M, output free
the model cannot invent page text, amounts, or labels outside the schema
one garbled page degrades one decision, flagged by needs_ocr
nine pages of noise bury the one clause that matters
Production code: the packet pipeline in TypeScript and Python
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
// Policy lives in application code — the model only reports facts.
const AUTO_FILE_CONFIDENCE = 0.85; // >= 0.85 -> file automatically
const REVIEW_MIN_CONFIDENCE = 0.6; // 0.60-0.85 -> review, below -> human
const QUESTIONS = {
is_single_document: {
type: "noul",
instructions:
"Does this page start a new document, or continue the document on the previous page?",
},
needs_ocr: {
type: "noul",
instructions:
"Is this page's text missing or too garbled to decide on, such that OCR must run first?",
},
doc_type: {
type: "choice",
instructions: "Classify this document into exactly one type",
criteria: {
invoice: "Itemized billing: line items, quantities, totals, tax, payment terms",
contract: "Binding agreement: parties, clauses, term dates, signature blocks",
receipt: "Proof of completed payment: paid marker, transaction reference, nothing due",
letter: "Human correspondence: salutation, body prose, signature block",
other: "None of the above matches with confidence",
},
},
} as const;
export type DocLane = "auto_file" | "review" | "human_review";
export async function evaluatePage(page: {
text: string;
page_index: number;
source: string;
has_ocr: boolean;
}) {
const response = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state: page, questions: QUESTIONS }),
});
if (!response.ok) throw new Error("Jev evaluate failed");
const data = await response.json();
const { is_single_document, needs_ocr, doc_type } = data;
const lane: DocLane =
needs_ocr.answer || doc_type.answer === "other"
? "human_review"
: doc_type.confidence >= AUTO_FILE_CONFIDENCE
? "auto_file"
: doc_type.confidence >= REVIEW_MIN_CONFIDENCE
? "review"
: "human_review";
return {
page_index: page.page_index,
starts_document: is_single_document.answer,
boundary_confidence: is_single_document.confidence,
needs_ocr: needs_ocr.answer,
doc_type: doc_type.answer,
doc_type_confidence: doc_type.confidence,
lane,
};
}
// Boundary and type answers come back in the same call: pages where
// starts_document = true open a new segment; the rest accumulate.
export function splitPacket(pages: Awaited<ReturnType<typeof evaluatePage>>[]) {
const segments: (typeof pages)[] = [];
for (const page of pages) {
if (page.starts_document || segments.length === 0) segments.push([page]);
else segments[segments.length - 1].push(page);
}
return segments;
}
// Pages are independent — fan them out concurrently. A 9-page packet
// stays under ~500ms wall time, input tokens only, 0% schema errors.
const results = await Promise.all(packetPages.map(evaluatePage));
const segments = splitPacket(results);
const reviewQueue = results.filter((r) => r.lane !== "auto_file");import requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_FILE_CONFIDENCE = 0.85 # >= 0.85 -> file automatically
REVIEW_MIN_CONFIDENCE = 0.60 # 0.60-0.85 -> review queue, below -> human
QUESTIONS = {
"is_single_document": {
"type": "noul",
"instructions": "Does this page start a new document, or continue the document on the previous page?",
},
"needs_ocr": {
"type": "noul",
"instructions": "Is this page's text missing or too garbled to decide on, such that OCR must run first?",
},
"doc_type": {
"type": "choice",
"instructions": "Classify this document into exactly one type",
"criteria": {
"invoice": "Itemized billing: line items, quantities, totals, tax, payment terms",
"contract": "Binding agreement: parties, clauses, term dates, signature blocks",
"receipt": "Proof of completed payment: paid marker, transaction reference, nothing due",
"letter": "Human correspondence: salutation, body prose, signature block",
"other": "None of the above matches with confidence",
},
},
}
def evaluate_page(page: dict) -> dict:
resp = requests.post(
JEV_ENDPOINT,
json={"state": page, "questions": QUESTIONS},
timeout=5,
)
resp.raise_for_status()
data = resp.json()
doc = data["doc_type"]
if data["needs_ocr"]["answer"] or doc["answer"] == "other":
lane = "human_review"
elif doc["confidence"] >= AUTO_FILE_CONFIDENCE:
lane = "auto_file"
elif doc["confidence"] >= REVIEW_MIN_CONFIDENCE:
lane = "review"
else:
lane = "human_review"
return {
"page_index": page["page_index"],
"starts_document": data["is_single_document"]["answer"],
"doc_type": doc["answer"],
"confidence": doc["confidence"],
"lane": lane,
}
def classify_packet(pages: list[dict]) -> list[list[dict]]:
# Pages are independent — fan out concurrently in production.
results = [evaluate_page(p) for p in pages]
segments: list[list[dict]] = []
for page in results:
if page["starts_document"] or not segments:
segments.append([page])
else:
segments[-1].append(page)
return segments
def scan_inbound_queue(packets: list[list[dict]]) -> dict:
# ~100-500ms per page, input-only pricing at $0.042/M: a
# 10,000-packet back-catalog scan costs about a dollar, no GPUs.
lanes: dict[str, int] = {}
for packet in packets:
for segment in classify_packet(packet):
lane = segment[0]["lane"]
lanes[lane] = lanes.get(lane, 0) + 1
return lanesDocument classification FAQ
How is document classification different from text classification?
Text classification labels one self-contained text — one ticket, one review, one message. Document classification faces packets: an email with attachments, a scanner batch that mixed eight invoice pages with a contract, a PDF export that concatenated unrelated documents. The two-stage pipeline adds the step text classification never needs — deciding where one document ends and the next begins with a per-page Noul — and only then classifies each segment with a typed Choice. The single-text general case is covered in the text classification API guide.
Can Jev read scanned PDFs, images, or handwriting?
No — and being upfront about this is the difference between a demo and a pipeline. Jev's state is text: there is no OCR, no layout analysis, no vision inside the model. Scanned packets run OCR and extraction upstream and pass the resulting text in. What Jev adds is a needs_ocr Noul that flags pages whose text is missing or too garbled to decide on, so they land in a repair queue instead of getting a confident wrong label. Garbled OCR caps the accuracy of every downstream classifier — typed or generative.
How does the packet splitting actually work?
Each page gets an is_single_document Noul: "does this page start a new document, or continue the previous one?" Pages where the answer is true (always the first page) open a segment; the other pages accumulate into the open segment. The calibrated confidence matters more than the boolean: a boundary at 0.55 confidence is a contested boundary, and contested boundaries — like every other low-confidence call — can drop into the review queue instead of silently separating a contract from its invoice.
What confidence thresholds should I use?
Start with the same three lanes as the confidence-gated fallback chain: ≥ 0.85 auto-files, 0.60–0.85 goes to a review queue, and below 0.60 — plus every "other" label and every needs_ocr page — escalates to a human. Then derive your own cutoffs from labeled samples: plot predicted confidence against ops verdicts on real packets and place each cutoff where a mis-filed document is cheap to catch. Shadow-mode the policy before enforcement and re-validate quarterly — new vendors mean new layouts mean distribution shift.
What is DocJev — and is it official?
DocJev is an open-source library by Jerry Liu, the founder of LlamaIndex, that classifies and splits complex document packets — the same two-stage pattern this page implements. It is featured on the madewithjev showcase under Documents & OCR, with the author reporting ~139ms per packet and 40/40 accuracy on his own benchmark (author-reported, not an independent audit). DocJev is a third-party ecosystem project, not an official TypeSafe AI product — and neither is madewithjev.
What does classifying a large batch cost?
Billing is input tokens only at $0.042/M and output is free. A 9-page packet at roughly 300 tokens per page is ~2,700 input tokens ≈ $0.0001, so a 10,000-packet back-catalog scan lands near a dollar with no GPUs. A published ecosystem datapoint: 1kpapers classified 1,018 AI research papers for $0.08 total at ~256ms median latency. Generative pipelines pay a comparable input bill plus output tokens for every verdict — and again for every retry.