Jev recipe / finance ops
Expense Categorization with Jev: Typed Bank Statement Pipeline
Turn raw card and bank statement lines into typed financial decisions: a category Choice, a deductibility Score, and a needs_review Noul — one evaluate call per statement, with a 0.85 confidence gate splitting auto-posting from human review.
Build the expense pipeline in six steps
{
"category": {
"type": "choice",
"instructions": "For each line, classify the expense into exactly one category",
"criteria": {
"software_saas": "SaaS subscriptions, cloud hosting, digital tools and licenses",
"travel": "Flights, hotels, ground transport, per-diem travel costs",
"meals": "Restaurants, coffee, client dinners and team lunches",
"office_supplies": "Stationery, equipment, furniture, warehouse purchases",
"professional_services": "Legal, accounting, consulting, contractors, agencies",
"other": "None of the above matches with confidence — escalate instead of forcing a label"
}
},
"deductibility": {
"type": "score",
"instructions": "For each line, rate tax deductibility, 1 (personal, not deductible) to 4 (fully deductible business expense)"
},
"needs_review": {
"type": "noul",
"instructions": "For each line, is it too ambiguous to categorize reliably — should a human bookkeeper confirm before posting?"
}
}Batch expense categorization simulator
Pick a statement and run the pipeline: a category Choice, a deductibility Score, and a needs_review Noul per line, with the 0.85 gate splitting auto-post, review queue, and bookkeeper lanes (front-end simulation only, no API calls):
AWS EMEA SARL
$184.20UBER *TRIP 88112
$23.75OFFICE DEPOT #221
$96.40SQ *BLUE BOTTLE COFFEE
$18.50GRANITE LEGAL PLLC
$2,400.00MERCH PAYMENT 8842 LLC
$312.88ZOOM.US
$15.99Press "Run batch categorization" to see per-line categories, deductibility scores, and the confidence split.
Expense categorization: Jev typed contract vs keyword rules and generative LLM
the next call already uses the new taxonomy; no retraining, no regex rewrite
0.85 behaves like an 85% accuracy contract, so thresholds gate auto-posting directly
a keyword either matches or it does not. Generative: self-reported, uncalibrated without extra logprob plumbing
unresolved lines become a review queue instead of a forced guess
30 lines ride one evaluate call in a single forward pass
a 30-line statement takes 45–90s
a 30-line statement ≈ $0.00007; 10,000 statements a month cost under a dollar
every misfile costs human minutes. Generative: output tokens billed per line, retries re-bill
the same line always lands in the same category with the same confidence, so books reconcile run to run
a misfile is explainable only by tracing regex lists. Generative: sampling variance flips borderline lines between runs
Production code: the expense pipeline in TypeScript and Python
import json
import os
import time
import requests
# Baseline: keyword rules first, a generative LLM for everything the
# rules cannot match. Two systems, two failure modes.
KEYWORD_RULES = {
"software_saas": ["AWS", "NOTION", "ZOOM.US", "ADOBE"],
"travel": ["DELTA AIR", "UBER", "LYFT", "MARRIOTT"],
"meals": ["RESTAURANT", "COFFEE", "SUSHI", "CAFE"],
"office_supplies": ["OFFICE DEPOT", "STAPLES", "AMZN MKTP"],
}
CATEGORIES = set(KEYWORD_RULES) | {"professional_services", "other"}
def rule_category(line: str) -> str | None:
upper = line.upper()
for category, keywords in KEYWORD_RULES.items():
if any(keyword in upper for keyword in keywords):
return category
return None # "SQ *BLUE BOTTLE", "MERCH PAYMENT 8842 LLC" match nothing
def generative_category(line: str) -> dict:
# 1.5-3s per line, output tokens billed per attempt, and json.loads
# accepts a hallucinated category as happily as a real one.
for attempt in range(3):
resp = requests.post(
"https://api.openai.com/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
json={
"model": "gpt-4o-mini",
"response_format": {"type": "json_object"},
"messages": [{
"role": "user",
"content": f'Categorize this expense line as one of {", ".join(sorted(CATEGORIES))}. '
f'Reply as JSON: {{"category": "..."}}\nLine: {line}',
}],
},
timeout=30,
)
try:
category = json.loads(
resp.json()["choices"][0]["message"]["content"]
)["category"]
if category in CATEGORIES: # value check json_mode does not give you
return {"category": category, "confidence": None}
except (KeyError, json.JSONDecodeError):
time.sleep(2 ** attempt) # every retry re-bills the generation
return {"category": "other", "confidence": None} # silent failure
def categorize_line(line: str) -> dict:
category = rule_category(line)
if category is not None:
return {"category": category, "confidence": None} # rules carry no confidence
return generative_category(line)import requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_POST_CONFIDENCE = 0.85 # >= 0.85 -> auto-post
REVIEW_MIN_CONFIDENCE = 0.60 # 0.60-0.85 -> review queue, below -> human
QUESTIONS = {
"category": {
"type": "choice",
"instructions": "For each line, classify the expense into exactly one category",
"criteria": {
"software_saas": "SaaS subscriptions, cloud hosting, digital tools",
"travel": "Flights, hotels, ground transport",
"meals": "Restaurants, coffee, client dinners",
"office_supplies": "Stationery, equipment, warehouse purchases",
"professional_services": "Legal, accounting, consulting, contractors",
"other": "None of the above matches with confidence",
},
},
"deductibility": {
"type": "score",
"instructions": "For each line, rate tax deductibility, 1 (personal) to 4 (fully deductible)",
},
"needs_review": {
"type": "noul",
"instructions": "For each line, is it too ambiguous to post without a bookkeeper confirming first?",
},
}
def categorize_statement(lines: list[dict]) -> list[dict]:
# The whole statement rides as one state — three typed questions
# answered per line in a single ~70-100ms forward pass.
resp = requests.post(
JEV_ENDPOINT,
json={"state": {"lines": lines}, "questions": QUESTIONS},
timeout=5,
)
resp.raise_for_status()
data = resp.json()
results = []
for i, line in enumerate(lines):
category = data["category"][i]
flagged = data["needs_review"][i]["answer"]
if flagged or category["answer"] == "other":
lane = "human_review"
elif category["confidence"] >= AUTO_POST_CONFIDENCE:
lane = "auto_post"
elif category["confidence"] >= REVIEW_MIN_CONFIDENCE:
lane = "review"
else:
lane = "human_review"
results.append({
"line_index": line["line_index"],
"category": category["answer"],
"confidence": category["confidence"],
"deductibility": data["deductibility"][i]["answer"],
"lane": lane,
})
return results
def sweep_statements(statements: list[list[dict]]) -> dict:
# Input-only pricing at $0.042/M: a 30-line statement is ~1,500
# tokens ≈ $0.00007, so 10,000 statements a month cost under a dollar.
lanes: dict[str, int] = {"auto_post": 0, "review": 0, "human_review": 0}
for lines in statements:
for row in categorize_statement(lines):
lanes[row["lane"]] += 1
return lanesimport { z } from "zod";
import OpenAI from "openai";
// Baseline: keyword rules first, structured-output LLM for the rest.
const KEYWORD_RULES: Record<string, string[]> = {
software_saas: ["AWS", "NOTION", "ZOOM.US", "ADOBE"],
travel: ["DELTA AIR", "UBER", "LYFT", "MARRIOTT"],
meals: ["RESTAURANT", "COFFEE", "SUSHI", "CAFE"],
office_supplies: ["OFFICE DEPOT", "STAPLES", "AMZN MKTP"],
};
const CATEGORIES = [
"software_saas",
"travel",
"meals",
"office_supplies",
"professional_services",
"other",
] as const;
const Decision = z.object({ category: z.enum(CATEGORIES) });
const client = new OpenAI();
function ruleCategory(line: string): string | null {
const upper = line.toUpperCase();
for (const [category, keywords] of Object.entries(KEYWORD_RULES)) {
if (keywords.some((k) => upper.includes(k))) return category;
}
return null; // "MERCH PAYMENT 8842 LLC" matches nothing
}
export async function categorizeLine(line: string) {
const category = ruleCategory(line);
if (category) return { category, confidence: null }; // rules carry no confidence
for (let attempt = 0; attempt < 3; attempt++) {
try {
const resp = await client.chat.completions.create({
model: "gpt-4o-mini",
response_format: { type: "json_object" },
messages: [
{
role: "user",
content: `Categorize this expense line as one of ${CATEGORIES.join(", ")}. Reply as JSON: {"category": "..."}\nLine: ${line}`,
},
],
});
// The shape is guaranteed; whether the value is right is not —
// and there is no confidence number to gate automation on.
return {
category: Decision.parse(
JSON.parse(resp.choices[0].message.content!),
).category,
confidence: null,
};
} catch {
await new Promise((r) => setTimeout(r, 2 ** attempt * 1000));
}
}
return { category: "other" as const, confidence: null };
}const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
// Policy lives in application code — the model only reports facts.
const AUTO_POST_CONFIDENCE = 0.85; // >= 0.85 -> auto-post
const REVIEW_MIN_CONFIDENCE = 0.6; // 0.60-0.85 -> review queue
type StatementLine = {
line_index: number;
merchant: string;
amount: number;
date: string;
memo?: string;
};
const QUESTIONS = {
category: {
type: "choice",
instructions:
"For each line, classify the expense into exactly one category",
criteria: {
software_saas: "SaaS subscriptions, cloud hosting, digital tools",
travel: "Flights, hotels, ground transport",
meals: "Restaurants, coffee, client dinners",
office_supplies: "Stationery, equipment, warehouse purchases",
professional_services: "Legal, accounting, consulting, contractors",
other: "None of the above matches with confidence",
},
},
deductibility: {
type: "score",
instructions:
"For each line, rate tax deductibility, 1 (personal) to 4 (fully deductible)",
},
needs_review: {
type: "noul",
instructions:
"For each line, is it too ambiguous to post without a bookkeeper confirming first?",
},
} as const;
export type ExpenseLane = "auto_post" | "review" | "human_review";
export async function categorizeStatement(lines: StatementLine[]) {
const response = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state: { lines }, questions: QUESTIONS }),
});
if (!response.ok) throw new Error("Jev evaluate failed");
const data = await response.json();
// Answers come back per line, aligned with state.lines order.
return lines.map((line, i) => {
const category = data.category[i] as {
answer: string;
confidence: number;
};
const flagged = data.needs_review[i].answer as boolean;
const lane: ExpenseLane =
flagged || category.answer === "other"
? "human_review"
: category.confidence >= AUTO_POST_CONFIDENCE
? "auto_post"
: category.confidence >= REVIEW_MIN_CONFIDENCE
? "review"
: "human_review";
return {
...line,
category: category.answer,
confidence: category.confidence,
deductibility: data.deductibility[i].answer as number,
lane,
};
});
}
// Amounts, tax, and per-category rollups stay in your code — the model
// never does arithmetic. A 30-line statement ≈ 1,500 input tokens,
// ~70-100ms wall time, $0 output tokens.Expense categorization FAQ
What is expense categorization?
Expense categorization assigns every spending line — a card charge, a bank statement entry, a receipt — to a category in your bookkeeping: software, travel, meals, office supplies, professional services. Accountants have done it by hand for centuries; the automation question is whether software can read the free-text merchant descriptor and make the same call. This page treats each line as a typed financial decision: a Choice over your taxonomy, a deductibility Score, and a needs_review flag, with calibrated confidence deciding what posts automatically and what a human confirms.
How do you categorize bank statements automatically?
Export the statement, parse it into line objects (merchant descriptor, amount, date, memo), and send the lines as state to one Jev evaluate call — the recipe above batches 30 demo transactions into a single request. Three typed questions come back per line: the category Choice, the deductibility Score, and a needs_review Noul. Your code applies the 0.85 gate: high-confidence lines post to the ledger, the rest land in a review queue. Parsing the export is your side of the boundary — Jev reads the text of the lines, not the PDF layout; that upstream split-then-classify problem is the document classification recipe.
Can AI categorize expenses accurately?
The honest answer is per-line, with a gate. Jev confidence is RLCD-calibrated, so a 0.85 threshold behaves like an 85% accuracy contract: at that gate roughly 85 of 100 auto-posted lines are correctly categorized — and the pipeline is built so the other 15 never post silently. Lines in the 0.60–0.85 band and every "other" or needs_review line go to a bookkeeper, whose corrections become labeled samples. Accuracy improves where it matters: not because the model retrained, but because the review queue keeps misfiles from ever reaching the books.
How is this different from generic LLM classification?
Three contract differences. First, bounded output: a generative classifier can answer "meals, probably — the vendor looks like a ramen chain" in prose; Jev returns one value from your taxonomy with a calibrated probability and nothing else. Second, the confidence number means something: RLCD calibration makes 0.85 an accuracy contract, while token logprobs measure typing confidence, not decision accuracy. Third, batching: one evaluate call answers all three questions for the whole statement at ~70–100ms with input-only billing — no output tokens per verdict, no retries. For the single-text REST and CLI general case, see the text classification API guide.
What categories should I use?
Start from the chart of accounts your bookkeeper already posts to, not from a model's idea of categories. Keep the axis small — five to eight categories plus a designed "other" — because every extra label dilutes probability mass, and each category should map to exactly one ledger account. Write criteria in the vocabulary your reviewers use ("client dinners", "cloud hosting"), version them in code like any other policy, and treat a merchant that keeps landing in "other" as a signal that a category or its criteria needs work.
What happens with unknown merchants?
Unknown is a designed outcome, not an error. An unparseable descriptor such as "MERCH PAYMENT 8842 LLC" gets the other label or a needs_review flag, drops below the 0.85 gate, and lands in the human confirmation queue — the human-in-the-loop lane that keeps one weird line from posting to the wrong account. Reviewers resolve it once, the correction joins the labeled set, and if the same unknown shape keeps appearing, that is your cue to add a category or sharpen its criteria.