Guides / illustrated walkthrough

Jev + Codex: Install the TypeSafe Skill and Triage a Real Gmail Inbox

A Chinese-language walkthrough of wiring Jev into OpenAI Codex via the typesafe-ai/skills repo: Playground drills on the three question types, the TYPESAFE_API_KEY environment variable, and a live run that priority-sorts 30 Gmail messages in 1.19 seconds of API time.

Quick takeaway

This is the missing "wire it into my agent" tutorial, originally in Chinese: after a Playground crash course on Choice, Noul, and Score (an interview invitation classified as a work email at 100%, flagged need-reply at 98%, and scored high priority), the creator installs the official typesafe-ai/skills repo into OpenAI Codex by simply asking Codex to do it. The skill is documentation, not a bundled model — the agent still needs TYPESAFE_API_KEY in the environment, set with one echo >> ~/.zshrc line. The payoff is a real inbox run: Codex reads the 30 most recent Gmail messages, and Jev sorts them into priority tiers — newsletters and verification codes at the bottom, a single Google security alert on top at 98%. The Jev API finished its share in 1.19 seconds; the whole task took 1 minute 40 seconds, almost all of it Codex thinking. The skill detects and registers itself for Codex, Claude Code, OpenCode, Hermes, Gemini CLI, GitHub Copilot, and Antigravity from one install.

Video source

Jason's Efficiency Workshop (杰森的效率工坊)

7:10SxJqfZmNaOA

Step-by-step walkthrough

  1. 1

    What a decision model answers: probabilities, not prose

    The video opens with the contrast that defines Jev: ask "will this stock go up or down" and a chat model writes an essay while Jev returns 72% up, 28% down — the driving-a-car-at-a-red-light argument for why A-or-B decisions should not start with 300 words of analysis. The first live demo analyzes a WhatsApp message from the creator's partner — "算了,你忙吧,我不打扰你了" ("fine, you're busy, I won't bother you") — and the response card reads like a relationship triage report: an 82% probability she does NOT want to end the conversation, "feels ignored" leading the intent breakdown at 76%, and a suggested next action: acknowledge you were not really listening before replying. A follow-up ("have I gained weight?") returns the equally honest verdict that no objective assessment was requested. The point of the jokes: intent, priority, and yes/no gates are decision problems, and Jev speaks them natively.

    WhatsApp message reading 算了你忙吧我不打扰你了 answered by a Jev AI card scoring 82 percent probability that she does not want to end the chat with feeling ignored leading the intent breakdown at 76 percent
    One chat message in, a probability report out: 82% "not actually ending the conversation", next action suggested.Watch at 0:42
  2. 2

    From signup to Playground: the console in one look

    Everything happens on typesafe.ai: register an account, open the API Console from the top-right corner, and the left navigation gives you the whole surface — Playground for experiments, Usage for consumption, API Keys for credentials, Documentation for the API contract. The home page doubles as a syllabus: the Learn to TypeSafe cards walk through NOUL ("Is hotdog a sandwich?"), CHOICE ("What color is the sky?"), and SCORE rating lessons, and a Quickstart panel already shows the agent-skill install commands this page comes back to later. The Playground input box takes a JSON state — the video frames it as "a prompt, but structured" — and the Questions panel below it is where typed questions get attached.

    TypeSafe API console home with the left navigation showing Home, Playground, Usage, API Keys and Documentation plus the Learn to TypeSafe lesson cards and a Quickstart agent setup panel
    The API Console: Playground to test, API Keys to provision, Documentation for the contract — the whole integration surface is four nav items.Watch at 2:10
  3. 3

    The three question types: Score, Choice, Noul

    Before wiring any agent, the video drills the primitives on the interview-invitation email it pastes into the state field. Choice offers named options and the model picks one: email_category returns work with a 100% confidence bar and flat zeros for the alternatives. Noul is a true/false gate: need_reply comes back true at 98% because the message explicitly asks the recipient to confirm attendance. Score attaches a rating scale — the creator defines low, medium, and high — and the response distributes probability across the scale instead of picking a bare label. All three can ride in one request, which is exactly what the inbox experiment exploits later.

    TypeSafe Playground example requests panel listing the three question types with lesson cards asking Is hotdog a sandwich, What color is the sky and Can monkeys create art
    The Playground's own lesson cards: Noul judges truth, Choice picks an option, Score grades on your scale.Watch at 2:28
  4. 4

    Choice in action: routing one email, 100% work

    The email_choice question defines the criteria inline — work emails that need a serious reply, personal mail, newsletters, promotions — and runs against the interview invitation. The answer panel returns email_category: work at a flat 100%, with every other option pinned at zero and the overall confidence stamped on the response. That is the shape a mail client, a CRM, or an agent needs: not a paragraph about the email, but a routing key with a confidence number attached. The creator's framing for builders: if you have a large volume of mail or records to pre-sort, this single call replaces the "read everything with a big model" step.

    TypeSafe Playground returning email_category work at 100 percent confidence for an interview invitation email with the option criteria listed under the Questions panel
    email_category → work, 100%. One typed call returns the routing key and its confidence.Watch at 3:05
  5. 5

    Noul in action: does this need a reply? 98% true

    The second question on the same email is a boolean: need_reply. The model answers true at 98% and shows its reasoning beside the answer — the message explicitly requests confirmation of attendance, which is about as unambiguous as a reply-requirement gets. This is the gate an automation actually branches on: below a threshold you auto-archive, above it you surface the mail to a human. The video's earlier WhatsApp demo was the same primitive aimed at messaage intent; here it is pointed at an inbox workflow.

    TypeSafe Playground need_reply question returning true at 98 percent probability for the same interview email with the reasoning summary shown beside the answer
    need_reply → true @ 98%: the boolean gate an automation can actually branch on.Watch at 3:25
  6. 6

    Score in action: grading urgency on your own scale

    The third primitive attaches a rubric. The creator defines a three-step priority scale and runs it against a cannot-login complaint email ("I keep trying the password and it still fails, please restore my account"). The response panel shows the priority question distributing probability across the three ratings with essentially all of it on the top rating and a 98%+ confidence stamp — exactly the kind of "deal with this now" signal a support queue needs before a human ever looks at it. Score is the primitive to reach for when "how much" matters, not just "which one".

    TypeSafe Playground priority score question placing about 100 percent probability on the top rating for a cannot-login complaint email with the three rating criteria listed beside it
    A complaint email scored against a low/medium/high rubric — the probability piles onto the top rating.Watch at 3:34
  7. 7

    All three at once: one chat log, three questions, one response

    To prove the primitives compose, the creator pastes a chat history into the state field and asks three differently-typed questions in a single request: is she upset with me (noul), how upset (score), and what does she actually want (choice over intent options). One response panel carries all three answers with their probabilities — very likely upset, feeling ignored, wants attentive listening. The video notes people have already wired this exact pattern into WeChat, and names the business versions that matter more: e-commerce customer service, complaint handling, and first-pass data screening. It also answers the obvious objection — yes, hand-writing JSON is unnatural, because these requests are meant to be emitted by agents, not typed by humans.

    TypeSafe Playground running three differently typed questions over one chat history and returning dissatisfaction level and underlying intent probabilities together in a single response panel
    Three question types, one request: upset? (noul), how much (score), what does she want (choice).Watch at 3:52
  8. 8

    Let Codex install the skill itself

    The official skill lives at github.com/typesafe-ai/skills, and the installation method the video actually demonstrates is the laziest one that works: paste the repo URL into Codex and ask it to install the skill. Codex clones the repo, detects every compatible agent on the machine, and reports back a completion card: global install finished in 1 minute 32 seconds, covering Codex, Claude Code, OpenCode, Hermes, Gemini CLI, GitHub Copilot, and Antigravity, with everything shared at ~/.agents/skills/typesafe-ai and each tool's entry verified readable. Restart Codex and the skill is live. Prefer doing it by hand? The docs show the equivalent one-liner for any agent: npx skills add typesafe-ai/skills --skill typesafe-ai -g, where -g means user-level rather than project-only.

    Codex install report for the typesafe-ai skill listing detected agents including Claude Code, Gemini CLI, GitHub Copilot and Antigravity with the shared path agents/skills/typesafe-ai and an elapsed time of 1 minute 32 seconds
    One install, eight agents: the skill registers itself for every compatible CLI it finds and shares one copy at ~/.agents/skills/typesafe-ai.Watch at 4:56
  9. 9

    The part people get wrong: the skill is not the model

    The video stops on the misconception before it costs anyone an evening: installing the skill teaches the agent HOW to call Jev — it ships no weights and includes no free quota. Real calls still need a credential, so the agent must find TYPESAFE_API_KEY in its environment. The docs page shown on screen gives the exact recipe: create a key on the API Keys page of the TypeSafe console, then append it to your shell profile with echo 'export TYPESAFE_API_KEY="我的key"' >> ~/.zshrc && source ~/.zshrc. From then on the agent decides on its own when a task needs a decision model — the skill's instructions tell it when to reach for Jev and how to build the evaluate call — and your only job is to keep the key valid. The Jev API is fully open, and OpenRouter works as an alternative channel if you would rather route through a gateway you already use.

    TypeSafe skill documentation showing the npx skills add typesafe-ai/skills install command and the echo export TYPESAFE_API_KEY line being appended to the zsh profile
    Two commands and you are done: install the skill for your agent, export TYPESAFE_API_KEY in your shell profile.Watch at 5:20
  10. 10

    The payoff: 30 real Gmail messages, priority-sorted

    The experiment that justifies the setup: Codex, already connected to Gmail, is asked to use the Jev skill to read the 30 most recent messages, sort them by priority, output a table, and report how long the work took. The inbox is realistically awful — newsletters, notifications, verification codes — and the table reflects it: row after row pinned at the bottom tiers, while one Google security alert is isolated at the top with a 98% score. Nothing was hand-labeled; every row's rating came from the same typed score question the Playground drills demonstrated. This is the "select A or B" workload the opening argued should never reach a long-context chat model.

    Codex output table sorting 30 Gmail messages into priority columns where a Google security alert ranks at 98 percent while newsletters and notification emails sit at the bottom tiers
    The real inbox run: 30 messages sorted, one security alert isolated at the top with 98%.Watch at 6:02
  11. 11

    The bill: 1.19 seconds of Jev, 100 seconds of agent

    The closing card separates the two clocks. The Jev API finished its share — the decision calls across 30 messages — in 1.19 seconds. The whole Codex task took 1 minute 40 seconds, and almost all of that was Codex itself: reading the mailbox, planning, and writing the table. Re-running the same 30 messages directly in the Playground returns answers near-instantly, confirming the model was never the bottleneck. The video's parting argument is the one to remember: attach a fast, cheap decision model to your agent workflow, and A-or-B questions stop costing big-model thinking time — whether that workflow is an AI customer-service line or a daily pass over your own inbox.

    Codex triage report noting the Jev API completed its 30 email decisions in 1.19 seconds while the whole task took 1 minute 40 seconds with the security alert row highlighted
    Jev: 1.19s for 30 decisions. The other ~99 seconds belonged to the agent, not the model.Watch at 6:42

Frequently asked questions

What is the Jev Codex skill?

It is the official agent-skill repository at github.com/typesafe-ai/skills: a set of instruction documents that teach a coding agent when a task calls for a typed decision and how to build the Jev evaluate call — which question types to use, how to structure the JSON state, and how to read the confidence numbers back. As the video stresses, it is documentation and tooling, not a bundled model: installing it changes what your agent knows, not what it can compute.

Does installing the skill make Jev free to use?

No — and the video calls this out explicitly. The skill only guides the agent; every real call still authenticates with your TYPESAFE_API_KEY and bills against your account. You create the key on the API Keys page of the TypeSafe API Console and expose it through the environment (echo 'export TYPESAFE_API_KEY="..."' >> ~/.zshrc && source ~/.zshrc). If you would rather not use the first-party endpoint, Jev is also reachable through platforms like OpenRouter.

How do I install the Jev skill for Codex?

Two routes, both shown in the video. The zero-effort route: paste https://github.com/typesafe-ai/skills into Codex and ask it to install the skill — it clones, registers itself for every compatible agent it detects (the video's machine covered Codex, Claude Code, OpenCode, Hermes, Gemini CLI, GitHub Copilot, and Antigravity in 1m32s), and you restart Codex. The manual route for any agent: npx skills add typesafe-ai/skills --skill typesafe-ai -g, where -g installs to the user-level directory (~/.agents/skills/typesafe-ai) instead of the current project.

What is the Gmail triage demo actually doing?

Codex reads the 30 most recent messages of a real inbox, then for each one asks Jev typed questions — what kind of mail is this, does it need a reply, how urgent is it on a low/medium/high scale — and assembles the answers into a priority table. In the video's inbox, newsletters and verification codes pile up in the bottom tiers while a single Google security alert is scored 98% and flagged high priority. The design generalizes to any A-or-B-heavy stream: customer-service tickets, form submissions, data records before analysis.

How fast was Jev in the email triage run?

The Jev API completed its decision calls across all 30 messages in 1.19 seconds. The end-to-end Codex task took 1 minute 40 seconds — nearly all of it the agent's own reading, planning, and table-writing, which is exactly why offloading decisions to a decision model pays. The same 30 messages pasted into the Playground answer near-instantly, confirming the model is never the slow part of this pipeline.

Codex or Claude Code for the Jev skill?

Both work from the same install — the skill registers itself for every compatible agent it detects, and the demo machine ended up covering eight CLIs at once. The division we use on this site: this page is the Codex-side walkthrough with the Gmail workflow, while our Jev + Claude Code guide covers the Claude Code side of the same integration (voice in, Jev decision, browser out). Install once, then use whichever agent is in front of you.

Related guides

More video walkthroughs