Jev recipe / finance ops

Expense Categorization with Jev: Typed Bank Statement Pipeline

Turn raw card and bank statement lines into typed financial decisions: a category Choice, a deductibility Score, and a needs_review Noul — one evaluate call per statement, with a 0.85 confidence gate splitting auto-posting from human review.

Build the expense pipeline in six steps

01Define the taxonomy before writing code. Start from the chart of accounts your bookkeeper already posts to — software_saas, travel, meals, office_supplies, professional_services, plus a designed other that escalates instead of forcing a label. Keep the axis small: every extra category dilutes probability mass, and every category should map to exactly one ledger account.
02Design the contract as three questions in one evaluate call: a category Choice with criteria per option, a deductibility Score (1 = personal, 4 = fully deductible) for tax treatment, and a needs_review Noul that asks whether the line is too ambiguous to post. The multi-question shape is the same typed contract the lead scoring recipe uses for sales priority — the model states facts, your code owns the policy.
03Own the gating policy in application code: category confidence ≥ 0.85 with needs_review = false auto-posts, the 0.60–0.85 band lands in the review queue, and below 0.60 — plus every "other" label — goes to a human bookkeeper. Derive cutoffs from labeled statements, never guess them; the 0.85 gate is an accuracy contract only because Jev confidence is RLCD-calibrated.
04Do the batch math before committing: state bills at $0.042 per million input tokens and output is free. A 30-line statement at roughly 50 tokens per line is ~1,500 input tokens ≈ $0.00007 per sweep, so categorizing 10,000 statements a month costs under a dollar — the full breakdown is on the Jev pricing page. The whole statement rides in one evaluate call.
05Close the loop with reconciliation: when a reviewer reclassifies a line, write the correction back to a labeled set and re-run it through the current schema in shadow mode. Corrections are how you find taxonomy gaps — a merchant that keeps landing in "other" means a category is missing or its criteria are too narrow, not noise to delete.
06Monitor distributions, not just accuracy: track the share of lines auto-posted, the "other" rate, and traffic in the contested 0.60–0.85 band. A new bank export format, a card rebrand, or a new spend category is distribution shift — the same drift the confidence-gated fallback chain recalibrates for. Re-validate thresholds quarterly.
schema / expense categorization contract
{
  "category": {
    "type": "choice",
    "instructions": "For each line, classify the expense into exactly one category",
    "criteria": {
      "software_saas": "SaaS subscriptions, cloud hosting, digital tools and licenses",
      "travel": "Flights, hotels, ground transport, per-diem travel costs",
      "meals": "Restaurants, coffee, client dinners and team lunches",
      "office_supplies": "Stationery, equipment, furniture, warehouse purchases",
      "professional_services": "Legal, accounting, consulting, contractors, agencies",
      "other": "None of the above matches with confidence — escalate instead of forcing a label"
    }
  },
  "deductibility": {
    "type": "score",
    "instructions": "For each line, rate tax deductibility, 1 (personal, not deductible) to 4 (fully deductible business expense)"
  },
  "needs_review": {
    "type": "noul",
    "instructions": "For each line, is it too ambiguous to categorize reliably — should a human bookkeeper confirm before posting?"
  }
}
Send a real statement line as state and inspect category, deductibility, and needs_review with calibrated confidence.
Interactive demo / expense categorization

Batch expense categorization simulator

Pick a statement and run the pipeline: a category Choice, a deductibility Score, and a needs_review Noul per line, with the 0.85 gate splitting auto-post, review queue, and bookkeeper lanes (front-end simulation only, no API calls):

category + deductibility + needs_review
7 lines
0103-01

AWS EMEA SARL

$184.20
0203-02

UBER *TRIP 88112

$23.75
0303-03

OFFICE DEPOT #221

$96.40
0403-04

SQ *BLUE BOTTLE COFFEE

$18.50
0503-05

GRANITE LEGAL PLLC

$2,400.00
0603-06

MERCH PAYMENT 8842 LLC

$312.88
0703-07

ZOOM.US

$15.99
source: corporate card / March
Jev decision output
~86ms / ≈$0.0000147

Press "Run batch categorization" to see per-line categories, deductibility scores, and the confidence split.

Policy in code: confidence >= 0.85 && needs_review = false -> auto-post · 0.60–0.85 -> review queue · < 0.60 or "other" or needs_review = true -> bookkeeper · amounts and tax stay in your code · demo data (simulated transactions)

Expense categorization: Jev typed contract vs keyword rules and generative LLM

METRIC
Jev
Keyword rules + generative LLM
Taxonomy changes
Edit versioned criteria in code

the next call already uses the new taxonomy; no retraining, no regex rewrite

Keyword rules rot as merchant descriptors change; a generative LLM carries the taxonomy in the prompt, where it drifts run to run
Confidence semantics
RLCD-calibrated probabilities

0.85 behaves like an 85% accuracy contract, so thresholds gate auto-posting directly

Rules: none

a keyword either matches or it does not. Generative: self-reported, uncalibrated without extra logprob plumbing

Unknown merchants
A designed "other" plus a needs_review Noul

unresolved lines become a review queue instead of a forced guess

Rules: unknown strings silently land in a default bucket. Generative: invents a plausible-looking category with no signal it guessed
Latency per statement
~70–100ms for the whole batch

30 lines ride one evaluate call in a single forward pass

Rules: instant but blind to free-text descriptors. Generative: 1.5–3s per line, token by token

a 30-line statement takes 45–90s

Batch cost
Input-only billing at $0.042/M, output free

a 30-line statement ≈ $0.00007; 10,000 statements a month cost under a dollar

Rules: free to run, expensive to correct

every misfile costs human minutes. Generative: output tokens billed per line, retries re-bill

Decision consistency
Deterministic

the same line always lands in the same category with the same confidence, so books reconcile run to run

Rules: deterministic but opaque

a misfile is explainable only by tracing regex lists. Generative: sampling variance flips borderline lines between runs

Production code: the expense pipeline in TypeScript and Python

python / baseline keyword rules + generative fallback
import json
import os
import time

import requests

# Baseline: keyword rules first, a generative LLM for everything the
# rules cannot match. Two systems, two failure modes.
KEYWORD_RULES = {
    "software_saas": ["AWS", "NOTION", "ZOOM.US", "ADOBE"],
    "travel": ["DELTA AIR", "UBER", "LYFT", "MARRIOTT"],
    "meals": ["RESTAURANT", "COFFEE", "SUSHI", "CAFE"],
    "office_supplies": ["OFFICE DEPOT", "STAPLES", "AMZN MKTP"],
}
CATEGORIES = set(KEYWORD_RULES) | {"professional_services", "other"}

def rule_category(line: str) -> str | None:
    upper = line.upper()
    for category, keywords in KEYWORD_RULES.items():
        if any(keyword in upper for keyword in keywords):
            return category
    return None  # "SQ *BLUE BOTTLE", "MERCH PAYMENT 8842 LLC" match nothing

def generative_category(line: str) -> dict:
    # 1.5-3s per line, output tokens billed per attempt, and json.loads
    # accepts a hallucinated category as happily as a real one.
    for attempt in range(3):
        resp = requests.post(
            "https://api.openai.com/v1/chat/completions",
            headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
            json={
                "model": "gpt-4o-mini",
                "response_format": {"type": "json_object"},
                "messages": [{
                    "role": "user",
                    "content": f'Categorize this expense line as one of {", ".join(sorted(CATEGORIES))}. '
                               f'Reply as JSON: {{"category": "..."}}\nLine: {line}',
                }],
            },
            timeout=30,
        )
        try:
            category = json.loads(
                resp.json()["choices"][0]["message"]["content"]
            )["category"]
            if category in CATEGORIES:  # value check json_mode does not give you
                return {"category": category, "confidence": None}
        except (KeyError, json.JSONDecodeError):
            time.sleep(2 ** attempt)  # every retry re-bills the generation
    return {"category": "other", "confidence": None}  # silent failure

def categorize_line(line: str) -> dict:
    category = rule_category(line)
    if category is not None:
        return {"category": category, "confidence": None}  # rules carry no confidence
    return generative_category(line)
python / jev batch statement categorization
import requests

JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_POST_CONFIDENCE = 0.85   # >= 0.85 -> auto-post
REVIEW_MIN_CONFIDENCE = 0.60  # 0.60-0.85 -> review queue, below -> human

QUESTIONS = {
    "category": {
        "type": "choice",
        "instructions": "For each line, classify the expense into exactly one category",
        "criteria": {
            "software_saas": "SaaS subscriptions, cloud hosting, digital tools",
            "travel": "Flights, hotels, ground transport",
            "meals": "Restaurants, coffee, client dinners",
            "office_supplies": "Stationery, equipment, warehouse purchases",
            "professional_services": "Legal, accounting, consulting, contractors",
            "other": "None of the above matches with confidence",
        },
    },
    "deductibility": {
        "type": "score",
        "instructions": "For each line, rate tax deductibility, 1 (personal) to 4 (fully deductible)",
    },
    "needs_review": {
        "type": "noul",
        "instructions": "For each line, is it too ambiguous to post without a bookkeeper confirming first?",
    },
}

def categorize_statement(lines: list[dict]) -> list[dict]:
    # The whole statement rides as one state — three typed questions
    # answered per line in a single ~70-100ms forward pass.
    resp = requests.post(
        JEV_ENDPOINT,
        json={"state": {"lines": lines}, "questions": QUESTIONS},
        timeout=5,
    )
    resp.raise_for_status()
    data = resp.json()

    results = []
    for i, line in enumerate(lines):
        category = data["category"][i]
        flagged = data["needs_review"][i]["answer"]
        if flagged or category["answer"] == "other":
            lane = "human_review"
        elif category["confidence"] >= AUTO_POST_CONFIDENCE:
            lane = "auto_post"
        elif category["confidence"] >= REVIEW_MIN_CONFIDENCE:
            lane = "review"
        else:
            lane = "human_review"
        results.append({
            "line_index": line["line_index"],
            "category": category["answer"],
            "confidence": category["confidence"],
            "deductibility": data["deductibility"][i]["answer"],
            "lane": lane,
        })
    return results

def sweep_statements(statements: list[list[dict]]) -> dict:
    # Input-only pricing at $0.042/M: a 30-line statement is ~1,500
    # tokens ≈ $0.00007, so 10,000 statements a month cost under a dollar.
    lanes: dict[str, int] = {"auto_post": 0, "review": 0, "human_review": 0}
    for lines in statements:
        for row in categorize_statement(lines):
            lanes[row["lane"]] += 1
    return lanes
typescript / baseline rules + structured-output retry
import { z } from "zod";
import OpenAI from "openai";

// Baseline: keyword rules first, structured-output LLM for the rest.
const KEYWORD_RULES: Record<string, string[]> = {
  software_saas: ["AWS", "NOTION", "ZOOM.US", "ADOBE"],
  travel: ["DELTA AIR", "UBER", "LYFT", "MARRIOTT"],
  meals: ["RESTAURANT", "COFFEE", "SUSHI", "CAFE"],
  office_supplies: ["OFFICE DEPOT", "STAPLES", "AMZN MKTP"],
};

const CATEGORIES = [
  "software_saas",
  "travel",
  "meals",
  "office_supplies",
  "professional_services",
  "other",
] as const;

const Decision = z.object({ category: z.enum(CATEGORIES) });
const client = new OpenAI();

function ruleCategory(line: string): string | null {
  const upper = line.toUpperCase();
  for (const [category, keywords] of Object.entries(KEYWORD_RULES)) {
    if (keywords.some((k) => upper.includes(k))) return category;
  }
  return null; // "MERCH PAYMENT 8842 LLC" matches nothing
}

export async function categorizeLine(line: string) {
  const category = ruleCategory(line);
  if (category) return { category, confidence: null }; // rules carry no confidence

  for (let attempt = 0; attempt < 3; attempt++) {
    try {
      const resp = await client.chat.completions.create({
        model: "gpt-4o-mini",
        response_format: { type: "json_object" },
        messages: [
          {
            role: "user",
            content: `Categorize this expense line as one of ${CATEGORIES.join(", ")}. Reply as JSON: {"category": "..."}\nLine: ${line}`,
          },
        ],
      });
      // The shape is guaranteed; whether the value is right is not —
      // and there is no confidence number to gate automation on.
      return {
        category: Decision.parse(
          JSON.parse(resp.choices[0].message.content!),
        ).category,
        confidence: null,
      };
    } catch {
      await new Promise((r) => setTimeout(r, 2 ** attempt * 1000));
    }
  }
  return { category: "other" as const, confidence: null };
}
typescript / jev typed statement pipeline
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";

// Policy lives in application code — the model only reports facts.
const AUTO_POST_CONFIDENCE = 0.85; // >= 0.85 -> auto-post
const REVIEW_MIN_CONFIDENCE = 0.6; // 0.60-0.85 -> review queue

type StatementLine = {
  line_index: number;
  merchant: string;
  amount: number;
  date: string;
  memo?: string;
};

const QUESTIONS = {
  category: {
    type: "choice",
    instructions:
      "For each line, classify the expense into exactly one category",
    criteria: {
      software_saas: "SaaS subscriptions, cloud hosting, digital tools",
      travel: "Flights, hotels, ground transport",
      meals: "Restaurants, coffee, client dinners",
      office_supplies: "Stationery, equipment, warehouse purchases",
      professional_services: "Legal, accounting, consulting, contractors",
      other: "None of the above matches with confidence",
    },
  },
  deductibility: {
    type: "score",
    instructions:
      "For each line, rate tax deductibility, 1 (personal) to 4 (fully deductible)",
  },
  needs_review: {
    type: "noul",
    instructions:
      "For each line, is it too ambiguous to post without a bookkeeper confirming first?",
  },
} as const;

export type ExpenseLane = "auto_post" | "review" | "human_review";

export async function categorizeStatement(lines: StatementLine[]) {
  const response = await fetch(JEV_ENDPOINT, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ state: { lines }, questions: QUESTIONS }),
  });
  if (!response.ok) throw new Error("Jev evaluate failed");
  const data = await response.json();

  // Answers come back per line, aligned with state.lines order.
  return lines.map((line, i) => {
    const category = data.category[i] as {
      answer: string;
      confidence: number;
    };
    const flagged = data.needs_review[i].answer as boolean;
    const lane: ExpenseLane =
      flagged || category.answer === "other"
        ? "human_review"
        : category.confidence >= AUTO_POST_CONFIDENCE
          ? "auto_post"
          : category.confidence >= REVIEW_MIN_CONFIDENCE
            ? "review"
            : "human_review";
    return {
      ...line,
      category: category.answer,
      confidence: category.confidence,
      deductibility: data.deductibility[i].answer as number,
      lane,
    };
  });
}

// Amounts, tax, and per-category rollups stay in your code — the model
// never does arithmetic. A 30-line statement ≈ 1,500 input tokens,
// ~70-100ms wall time, $0 output tokens.

Expense categorization FAQ

What is expense categorization?

Expense categorization assigns every spending line — a card charge, a bank statement entry, a receipt — to a category in your bookkeeping: software, travel, meals, office supplies, professional services. Accountants have done it by hand for centuries; the automation question is whether software can read the free-text merchant descriptor and make the same call. This page treats each line as a typed financial decision: a Choice over your taxonomy, a deductibility Score, and a needs_review flag, with calibrated confidence deciding what posts automatically and what a human confirms.

How do you categorize bank statements automatically?

Export the statement, parse it into line objects (merchant descriptor, amount, date, memo), and send the lines as state to one Jev evaluate call — the recipe above batches 30 demo transactions into a single request. Three typed questions come back per line: the category Choice, the deductibility Score, and a needs_review Noul. Your code applies the 0.85 gate: high-confidence lines post to the ledger, the rest land in a review queue. Parsing the export is your side of the boundary — Jev reads the text of the lines, not the PDF layout; that upstream split-then-classify problem is the document classification recipe.

Can AI categorize expenses accurately?

The honest answer is per-line, with a gate. Jev confidence is RLCD-calibrated, so a 0.85 threshold behaves like an 85% accuracy contract: at that gate roughly 85 of 100 auto-posted lines are correctly categorized — and the pipeline is built so the other 15 never post silently. Lines in the 0.60–0.85 band and every "other" or needs_review line go to a bookkeeper, whose corrections become labeled samples. Accuracy improves where it matters: not because the model retrained, but because the review queue keeps misfiles from ever reaching the books.

How is this different from generic LLM classification?

Three contract differences. First, bounded output: a generative classifier can answer "meals, probably — the vendor looks like a ramen chain" in prose; Jev returns one value from your taxonomy with a calibrated probability and nothing else. Second, the confidence number means something: RLCD calibration makes 0.85 an accuracy contract, while token logprobs measure typing confidence, not decision accuracy. Third, batching: one evaluate call answers all three questions for the whole statement at ~70–100ms with input-only billing — no output tokens per verdict, no retries. For the single-text REST and CLI general case, see the text classification API guide.

What categories should I use?

Start from the chart of accounts your bookkeeper already posts to, not from a model's idea of categories. Keep the axis small — five to eight categories plus a designed "other" — because every extra label dilutes probability mass, and each category should map to exactly one ledger account. Write criteria in the vocabulary your reviewers use ("client dinners", "cloud hosting"), version them in code like any other policy, and treat a merchant that keeps landing in "other" as a signal that a category or its criteria needs work.

What happens with unknown merchants?

Unknown is a designed outcome, not an error. An unparseable descriptor such as "MERCH PAYMENT 8842 LLC" gets the other label or a needs_review flag, drops below the 0.85 gate, and lands in the human confirmation queue — the human-in-the-loop lane that keeps one weird line from posting to the wrong account. Reviewers resolve it once, the correction joins the labeled set, and if the same unknown shape keeps appearing, that is your cue to add a category or sharpen its criteria.

Extend the finance pipeline