Guides / illustrated walkthrough

Image Decision Models for RPA: Watch Two Open Models Review an Unsigned W-8BEN With No OCR

Sam Witteveen’s RPA form-review demo turned into a step-by-step page: the flowchart framing (rectangles are steps, diamonds are decisions), the UiPath cautionary tale, the ImageJevBench leaderboard read frame by frame (Imajev-4B 76.39, Jev-Omni 73.10, NeoHorse Jev 4B 71.94), then the Form Decision Inspector running Imajev-4B and Jev-Omni side by side over an unsigned W-8BEN — unknown probabilities, a 90% auto-accept threshold, a true/false disagreement the confident model gets wrong, a multiple-choice rephrase both models agree on, a purely conditional follow-up email, and question engineering where adding a question is literally adding a text box.

Quick takeaway

The first non-text-input page in this guide series: two open image decision models — Imajev-4B (mohit67890/imajev-4b) and Jev-Omni (akhilaaa3/Jev-Omni) — review scanned forms inside a Form Decision Inspector app, and no OCR runs anywhere in the pipeline. Why it matters: RPA automated the rectangles of business processes (copy, click, file) but its flagship UiPath went from a ~$30-40B IPO valuation in 2021 (shares in the high $70s) to $13-14 per share and under $6.5B market cap, because per-task tuned deep learning never produced a universal classifier for the diamonds — the human judgment nodes. Image decision models attack exactly those diamonds: hand the model the raw image (screenshot, scan, photo) as state, ask typed questions (the app marks yes/no questions as noul and option questions as choice), and get a probability for every allowed option — including a trained unknown option, so the model either commits or abstains. The benchmark: ImageJevBench v0.1.3, 684 scored items (228 public, 456 sealed, 93 retired; the 333 fresh sealed items are disclosed as synthetic), 49 systems on the composite board. Top five, read off the zoomed frame: Imajev-4B 76.39, Jev-Omni 73.10, NeoHorse Jev 4B 71.94, Visual-Jev 4B Answer-SFT 69.81, JPT-4B 69.55 — the transcript’s garbled “Neo Horse” is NeoHorse Jev 4B. Imajev-4B’s origin is the lean-build story of the year: one person, 15 days, $1,200 of GPU rental, and a Reddit post by a refunds/support process consultant explaining that ambiguity made every decision node end on a human. The demo form: an unsigned W-8BEN with the printed name Oliver Bennett, the date 03-25-2026, a filled DOB field (11-16-1964), and an empty signature line highlighted in red. Run 1 (r20261002-195810-e908): 7 questions, 3.3 s total, Imajev-4B 3,159 ms vs Jev-Omni 1,072 ms on the page; Imajev’s form verdict is Return (failed: missing), Jev-Omni’s is Human (below 90% or abstained). The headline disagreement on signed: Imajev-4B True 91.5% (unknown 1.0%) passes the 90% auto-accept slider and is wrong; Jev-Omni False 73.2% sits below it, gets flagged Human, and is right — Auto is a confidence badge, not a truth badge. Both agree on dated (True 98.1% / 98.6%). Rephrase signed as the choice “What is on the signature line?” (handwritten signature / blank / typed name only) and both answer typed name only — 74.2% vs 75.1%, agree marker, Jev-Omni’s P(handwritten signature) down at 3.92% — both Human-flagged under the now-89% threshold, but correct: only a printed name. The follow-up email is pure conditional logic — “Action needed: your tax form… could you please sign the form in the signature field” — keyed off Imajev’s missing row (page 1, signature, 92.5%, Auto); no LLM, no OCR. Question engineering: adding a date-of-birth option to missing is adding one text box, no retraining. Re-run (3.4 s): Imajev-4B unbothered — still signature 94.5% (nothing missing 3.1%, date of birth 0.6%) — while Jev-Omni breaks: nothing missing 45.3% with date of birth second at 30.7%, and False 88.1% on the dob question Imajev answers True 99.4%. A second, signed form without a DOB shows both models confirming the handwritten signature and both catching the missing DOB at different confidences — model agreement is the real signal. The playbook: find which question shapes fit your purpose, then which model fits those shapes; set per-question confidence thresholds; route below-threshold answers to another model or a human. Use cases: RPA, rapid image classification, large-scale image processing.

Video source

Sam Witteveen

16:22L8YxigQoLaM

Step-by-step walkthrough

  1. 1

    Draw any process as a flowchart: rectangles are steps, diamonds are decisions

    Every business process drawn as a flowchart produces two shapes. Rectangles are steps — copy this, click that, file it — and diamonds are decisions: does this look right? Is this an invoice? Is this a payment? For years RPA tools have automated the rectangles: robotic process automation, software bots clicking screens and copying data between systems just like a human would, following fixed rules. The diamonds kept requiring a human, because a rule cannot look at a page and tell whether there is a signature on it. Deep learning made progress, but each company and each use case needed its own tuned model — and that is precisely the gap the rest of the video attacks with image decision models.

    A business process flowchart labels three hatched rectangles as steps with three question-mark diamonds between them, the shape vocabulary the video uses to separate bot work from human decisions.
    The whole thesis in one line of shapes: rectangles are automatable, diamonds are the humans.Watch at 0:08
  2. 2

    Meet the two open models: Imajev-4B and Jev-Omni, no OCR step

    The demo’s architecture fits in one diagram: PDF and JPG documents flow into an inspector, which passes each page to two small open models in parallel — Imajev-4B (Hugging Face mohit67890/imajev-4b) and Jev-Omni (akhilaaa3/Jev-Omni) — and collects yes/no answers. Both models rank near the top of the image Jev benchmark, and the handwritten note under the diagram is the headline: no OCR step. The raw image goes straight into the model; nothing recognizes text first. One timing note from the narrator: at recording, even Jev itself could not take images as input, so the image-decision axis belonged entirely to open models like these two.

    An OpenJev diagram routes PDF and JPG documents through an inspector box into Imajev-4B and Jev-Omni, ending at a yes-no diamond with a handwritten no OCR step note.
    Two open models behind one inspector: raw image in, yes/no out — and the OCR stage simply does not exist.Watch at 1:13
  3. 3

    The UiPath cautionary tale: from a ~$30-40B valuation to under $6.5B

    Why this matters commercially: RPA became a huge industry, and UiPath was its flagship — the narrator first encountered it in 2018 and found it genuinely impressive. By 2021 it was a public company at a ~$30-40B valuation with shares in the high $70s. The card then bends an orange line down to now: $13-14 per share, under a $6.5B market cap (the axis is stamped “not to scale”). Talking to people who worked with UiPath produced the diagnosis: not a bad product, but the deep learning of that era required significant tuning per task — there was no universal classifier that could hold high accuracy across many kinds of decision nodes. Image decision models are the candidate fix the rest of the video stress-tests.

    A UiPath timeline chart marks the 2021 IPO at a 30-40 billion dollar valuation with shares in the high 70s, then bends an orange line down to now at 13-14 dollars per share and under a 6.5 billion market cap.
    The RPA boom and bust in one chart: IPO at ~$30-40B, now under $6.5B — the diamonds were never automated.Watch at 3:24
  4. 4

    ImageJevBench v0.1.3: 684 scored items, 49 systems

    The benchmark card in full: ImageJevBench v0.1.3, “a held-out comparison of systems that make decisions from images, from interface targets to everyday scenes.” The frozen benchmark has 684 scored items — 228 public and 456 sealed — with 93 further items retired and not scored. A caveat box discloses its own methodology: the 333 fresh sealed items are the benchmark’s synthetic images and renders, easier for frontier API models than the older real-source items. The composite score spans 49 systems, pink bars are hosted APIs, and a system whose Cost or Calibration axis falls under the gate scores 0. This is the board the demo’s two models come from.

    The ImageJevBench v0.1.3 card describes 684 scored items split into 228 public and 456 sealed with 93 retired, warns that 333 fresh sealed items are synthetic, and opens its composite score of 49 systems with Imajev-4B at 76.39.
    The scoreboard’s own fine print: 684 scored items, 456 sealed, synthetic-item caveat disclosed, 49 systems ranked.Watch at 5:10
  5. 5

    Read the board: Imajev-4B 76.39 — and “Neo Horse” is NeoHorse Jev 4B

    Zoom into the top of the composite board and the rows resolve: 1. Imajev-4B 76.39 (the gold bar), 2. Jev-Omni 73.10, 3. NeoHorse Jev 4B 71.94, 4. Visual-Jev 4B Answer-SFT 69.81, 5. JPT-4B (kirp / llm2jev) 69.55. The auto-transcript’s garbled “Neo Horse” is NeoHorse Jev 4B at 71.94. The winner’s backstory is the lean-build story here: Imajev-4B was created by one person in 15 days, spending only $1,200 on GPU rental — and that person did not come from a research lab, which the next frame makes literal.

    A zoomed crop of the ImageJevBench top five rows reads Imajev-4B 76.39, Jev-Omni 73.10, NeoHorse Jev 4B 71.94, Visual-Jev 4B Answer-SFT 69.81 and JPT-4B 69.55.
    The composite top five: Imajev-4B 76.39, Jev-Omni 73.10, NeoHorse Jev 4B 71.94 — the transcript’s “Neo Horse,” settled frame by frame.Watch at 5:24
  6. 6

    Why the top model exists: a process consultant’s Reddit post

    The card traces Imajev-4B to its author — a process consultant — and to the Reddit post where he explained not just how the model was built but why. His background is refunds, products and customer support, and his point was that in those process maps almost every decision-making node ultimately depended on a human, because ambiguity is very difficult to handle with code alone. That is the demand side of this whole category: millions of flowchart diamonds where someone currently looks at a document and decides. The demo picks the board’s top two models — Imajev-4B and Jev-Omni — and points them at the most ordinary diamond of all: is the form signed?

    An Imajev-4B tag points to a person icon labeled process consultant and then to the Reddit logo, with a refunds tag hanging below the label.
    The author’s path: process consultant, refunds and support workflows, then a Reddit post on why every decision node ended on a human.Watch at 5:35
  7. 7

    Same questions as text Jev — the state is the raw image

    The mental model carries over from text decision models with one swap. A text Jev takes text as the state plus typed questions; an image Jev takes the image as (part of) the state and the same question shapes — signed? yes | no, date filled? yes | no, doc type? form | invoice. There is no OCR stage feeding the model a transcription; the model decides what is depicted and answers from pixels. The extra worth noticing: image Jev models are trained to also output a probability for the unknown option, so the model is either sure or says it does not know — which takes the guesswork out. The answers then drive the same smart if-conditions as text models, plus one upgrade: a confidence threshold below which the result is routed to another model for confirmation or to a human for review.

    A state box holds form_page.png beside a questions box listing signed, date filled and doc type options, both feeding an ImageJev decision diamond under a same questions note.
    Image as state, same typed questions: yes/no and choice shapes — with an unknown option the model is allowed to pick.Watch at 6:45
  8. 8

    The form: an unsigned W-8BEN with a printed name and a date

    The demo drags a US tax form PDF into Form Decision Inspector — “Inspect tax PDFs or JPGs. Run two image models, review disagreements.” The scan is a W-8BEN, Certificate of Foreign Status of Beneficial Owner: a Bristol BS1 9XA (United Kingdom) address, a filled date of birth field (11-16-1964), the printed name Oliver Bennett and the date 03-25-2026 in the Sign Here block — and an empty signature line above the printed name, which the viewer highlights in red. The question set covers exactly the diamonds a forms team cares about: is it signed, is the date next to the signature filled, is the mailing address present, is the date of birth filled, which required field is empty, what kind of form is it, how legible is the handwriting. DPI is adjustable, and both models run at once.

    The demo’s W-8BEN tax form with the Sign Here box highlighted in red shows an empty signature line above the printed name Oliver Bennett and the date 03-25-2026.
    The test document: everything filled except the one thing that matters — the signature line is empty.Watch at 8:28
  9. 9

    First run: 7 questions in 3.3 s — and the confident answer is the wrong one

    The results header reads total 3.3 s, 1 page, 7 questions; Imajev-4B spent 3,159 ms on the page against Jev-Omni’s 1,072 ms. On signed — “The form has a handwritten signature in the signature field” — Imajev-4B answers True 91.5% (false 8.5%, unknown 1.0%) and clears the auto-accept slider set at 90%, earning the green Auto badge; Jev-Omni answers False 73.2% (true 26.8%) and gets the yellow Human flag for sitting below the threshold. The form is not signed: the flagged answer is correct and the auto-accepted one is wrong — Auto means confident, not correct. The verdict row splits the same way: Return for Imajev-4B (failed: missing) versus Human for Jev-Omni (below 90% or abstained: signed, missing). On dated, both models agree — True 98.1% and 98.6% — and the date-of-birth row splits them again, one saying it is there and one saying it is not.

    Form Decision Inspector results for the signed question show Imajev-4B answering true at 91.5 percent with an Auto badge while Jev-Omni answers false at 73.2 percent flagged Human under a 90 percent auto-accept slider.
    The honesty frame: True 91.5% Auto is wrong, False 73.2% Human is right — confidence and correctness are different axes.Watch at 9:30
  10. 10

    Rephrase the diamond: true/false becomes a choice, and both models agree

    The fix is question engineering, not model swapping. Turn signed into a choice — “What is on the signature line?” with options handwritten signature, blank, and typed name only — and run again. Both models now return typed name only: Imajev-4B at 74.2% (blank 24.7%, handwritten signature 1.1%, unknown 0.2%), Jev-Omni at 75.1% with its handwritten-signature probability down at 3.92% per the inspector’s tooltip — and an agree marker sits between the cards. Both answers sit below the auto-accept threshold (now set at 89%) so both carry Human flags, but the answers match the ground truth: only a printed name, no handwritten signature. The narrator’s rule generalizes: on this task, this model is better at classification-shaped questions than true/false ones — find which question types fit your purpose, then which model fits those questions.

    Rephrased as what is on the signature line, both models return typed name only — Imajev-4B at 74.2 percent and Jev-Omni at 75.1 percent — with an agree marker between the two cards.
    Same diamond, better shape: typed name only 74.2% ≡ 75.1% — the rephrase turns a disagreement into agreement.Watch at 11:20
  11. 11

    The follow-up email: pure conditional logic — no LLM, no OCR

    The Follow-up email panel composes the chase-down message with zero generation: subject “Action needed: your tax form,” a toggle between Imajev-4B’s answers and Jev-Omni’s answers, Copy and Open in mail app buttons, and a template body — “Thanks for sending your tax form. Before we can process it, could you please: • Sign the form in the signature field. You can reply to this email with the corrected form attached.” No LLM writes a word of it; it is conditional logic keyed off the model answers, with the reasoning trace shown underneath (missing, page 1, signature, 92.5%, Auto). And as the narrator re-stresses, there is no optical character recognition anywhere in this loop — the image is transmitted and processed as an image, and the email just formats the decisions.

    The Follow-up email panel drafts Action needed: your tax form, telling the sender to sign the form in the signature field, justified by the missing row reading 92.5 percent Auto.
    Ask the sender to fix it: a template email triggered by one answer — the LLM-free half of the workflow.Watch at 12:10
  12. 12

    Add a question, add a text box — then re-test every model

    Adding a new class to the missing question is literally adding an option: date, address, nothing missing — plus a new date of birth. No retraining anywhere. On the re-run (total 3.4 s) Imajev-4B is unbothered: it still answers signature 94.5% for which required field is empty (nothing missing 3.1%, date of birth 0.6%). Jev-Omni breaks: nothing missing 45.3% with date of birth second at 30.7% — “nothing missing… maybe the date of birth is missing” — and on the dob question itself it answers False 88.1% where Imajev-4B answers True 99.4%. A second form — signed this time, but missing the date of birth — then shows both models confirming the handwritten signature and both catching the missing DOB with different confidence. That is the real lesson of running two models: agreement is the signal, disagreement is the routing condition, and every added question re-runs the fit test. Per-question confidence thresholds, multiple-choice questions, scored questions — all just more text boxes, tailored to the use case: RPA, rapid image classification, large-scale image processing.

    After a date of birth option is added to the missing question, Imajev-4B keeps answering signature at 94.5 percent while Jev-Omni drifts to nothing missing at 45.3 percent with date of birth second at 30.7 percent.
    One new option, no retraining: Imajev holds at signature 94.5%, Jev-Omni drifts to “nothing missing” 45.3% — re-test after every change.Watch at 13:03

Frequently asked questions

Is RPA still worth learning?

Yes — with clearer eyes about which half of the job is automatable. The UiPath arc on the video’s chart is real: a ~$30-40B IPO valuation in 2021 and shares in the high $70s, down to $13-14 per share and under a $6.5B market cap. That collapse is not proof automation failed; it is proof that per-task tuned deep learning could not deliver a universal classifier for decision nodes — the diamonds in a process map. The rectangles (copy, click, file) have been bot territory for a decade. The diamonds are the open half, and image decision models with confidence thresholds and human-review routing are the first credible tool aimed squarely at them.

What is an image decision model?

A Jev-style decision model whose state is a raw image — a screenshot, a scan, or a photo — instead of text. You hand it the image plus a set of typed questions (yes/no questions the app labels noul, option questions labeled choice), and it returns a probability for every allowed option of every question in one pass. No OCR runs before the model: it reads pixels directly. These models are also trained to output a probability for an unknown option, so instead of guessing on an ambiguous form they can abstain — and the answers plug into the same conditional logic as text decision models, with a confidence threshold deciding what goes to a human.

Why run two models side by side?

Because they fail differently, and agreement is the only cheap signal you get. In the demo’s first run, Imajev-4B said the unsigned W-8BEN was signed — True 91.5%, above the 90% auto-accept bar, and wrong — while Jev-Omni said False 73.2%, below the bar, flagged for human review, and right. On the what-is-missing question after a new option was added, Imajev-4B held at signature 94.5% while Jev-Omni drifted to nothing missing 45.3%. When the signature question was rephrased as a multiple choice, both answered typed name only (74.2% and 75.1%). One model is not better; each is better at different question shapes — so running both turns disagreement into a routing rule instead of a silent failure.

What is the unknown option actually worth?

It converts silent guesses into explicit abstentions. These models output a probability for unknown alongside every real option — 1.0% on Imajev-4B’s signature answer, 0.2% on the multiple-choice rephrase — so “I can’t tell” becomes a measurable outcome rather than a low-confidence guess you have to reverse-engineer. Combined with a confidence threshold (90% in the first run, 89% in later re-runs, adjustable per question), anything uncertain or below the bar routes to another model for confirmation or to a human. In a forms-review workflow that is exactly the behavior you want: confident answers automate, uncertain answers escalate.

What are Imajev-4B and Jev-Omni?

Two community open-weight image decision models on Hugging Face — mohit67890/imajev-4b and akhilaaa3/Jev-Omni — that sit first and second on ImageJevBench v0.1.3’s composite board (76.39 and 73.10 out of 49 systems, ahead of NeoHorse Jev 4B at 71.94). Imajev-4B is the notable story: built by one person in 15 days on $1,200 of GPU rental, by a business process consultant who posted on Reddit that ambiguity kept every decision node in his clients’ workflows on a human. Both models accept raw images with no OCR step. At the time of recording, Jev itself could not take image input, which is exactly why the image axis belonged to open models like these.

How is this different from text-based form filling?

The state is different, so the jobs do not overlap. Text form filling (see our Jev form filling guide) works on structured text you already have: the state is a string, the questions run over an API, and the model never sees a pixel. This page’s workflow starts where there is no text layer at all — scanned PDFs, phone photos, screenshots — and the model reads the raw image, so questions like “is there a handwritten signature” or “which required field is empty” are answerable without an OCR stage. In practice they compose: text decision models drive the structured filling and routing, image decision models verify what actually landed on the page — and both feed the same kind of conditional follow-up logic.

Related guides

More video walkthroughs