Jev Recipe / 厂商对比
Jev vs OpenAI Decision API(Luna):决策 API 正面对比
OpenAI 在 DevDay 用 GPT-6 Luna 的 Decision API 正面回应 Jev。这份对比看契约本身:预定义答案集 vs Choice/Score/Noul 三原语、150ms vs ~95ms、定价未公布 vs 公开透明,以及两种置信度各自值多少信任。
选定决策 API 之前的 6 项检查
{
"routing": {
"type": "choice",
"instructions": "Which queue should this ticket go to?",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Product malfunction or bug",
"sales": "Buying or upgrade question",
"abuse": "Safety or abuse report"
}
},
"needs_safety_escalation": {
"type": "noul",
"instructions": "Does this ticket require a safety escalation regardless of queue?"
},
"urgency": {
"type": "score",
"instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)"
}
}决策 API 契约对比:Jev 类型化原语 vs OpenAI Decision API(Luna)
代码对比:OpenAI decisions 端点 vs Jev 类型化调用
import os
import requests
# Limited-preview shape (announced 2026-09-29). Verify against the
# current reference at broad rollout - preview APIs move.
resp = requests.post(
"https://api.openai.com/v1/decisions",
headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
json={
"model": "gpt-6-luna",
"context": "Ticket: my invoice shows the same charge twice...",
"question": "Which queue should this ticket go to?",
"answers": ["billing", "technical", "sales", "abuse"],
},
timeout=5,
)
resp.raise_for_status()
decision = resp.json() # -> {"answer": "billing", "confidence": 0.91}
# One flat answer set per call: no per-option criteria, no second
# question sharing the call, no ordered score primitive.import requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_ROUTE_CONFIDENCE = 0.85
QUESTIONS = {
"routing": {
"type": "choice",
"instructions": "Which queue should this ticket go to?",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Product malfunction or bug",
"sales": "Buying or upgrade question",
"abuse": "Safety or abuse report",
},
},
"needs_safety_escalation": {
"type": "noul",
"instructions": "Does this ticket require a safety escalation regardless of queue?",
},
"urgency": {
"type": "score",
"instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
},
}
resp = requests.post(
JEV_ENDPOINT,
json={"state": {"ticket": "..."}, "questions": QUESTIONS},
timeout=5,
)
resp.raise_for_status()
data = resp.json()
route = data["routing"] # criteria-checked answer + calibrated confidence
if (
route["confidence"] >= AUTO_ROUTE_CONFIDENCE
and not data["needs_safety_escalation"]["answer"]
):
lane = f"auto:{route['answer']}" # ~70-100ms, input tokens only
else:
lane = "review" # 0.60-0.85 band or safety flagconst resp = await fetch("https://api.openai.com/v1/decisions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-6-luna",
context: "Ticket: my invoice shows the same charge twice...",
question: "Which queue should this ticket go to?",
answers: ["billing", "technical", "sales", "abuse"],
}),
});
// Limited-preview shape (announced 2026-09-29) - verify at broad rollout.
const decision = (await resp.json()) as {
answer: string;
confidence: number;
};const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_ROUTE_CONFIDENCE = 0.85;
const QUESTIONS = {
routing: {
type: "choice",
instructions: "Which queue should this ticket go to?",
criteria: {
billing: "Payment, invoice or refund issue",
technical: "Product malfunction or bug",
sales: "Buying or upgrade question",
abuse: "Safety or abuse report",
},
},
needs_safety_escalation: {
type: "noul",
instructions:
"Does this ticket require a safety escalation regardless of queue?",
},
urgency: {
type: "score",
instructions:
"Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
},
} as const;
const resp = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state: { ticket: "..." }, questions: QUESTIONS }),
});
const data = await resp.json();
const lane =
data.routing.confidence >= AUTO_ROUTE_CONFIDENCE &&
!data.needs_safety_escalation.answer
? `auto:${data.routing.answer}`
: "review";Jev vs OpenAI Decision API 常见问题
OpenAI 的 Decision API 是什么?
OpenAI 在 2026-09-29 DevDay 官宣的非聊天端点,构建于 GPT-6 家族中最小、最便宜的模型 Luna 之上。调用方提供问题、预定义答案集和上下文;它在大约 150ms 内返回答案集中的一个答案加一个置信度分数。首发为 limited preview,官方称「未来几天」扩大开放;单次调用定价在官宣时仍未公布。The New Stack 报道认为它是对 TypeSafe Jev 的直接回应。
Decision API 和 Jev 是同一个东西吗?
同一个品类,不同的契约。两者都非生成式:上下文 + 问题 + 固定答案集进,单一答案 + 置信度出——不生成任何文本。区别在契约深度:Jev 的 Choice 问题给每个选项带 criteria(策略留在可评审的代码里)、Score 覆盖有序决策、多问题共用一次调用、概率经 RLCD 校准、仅输入计费且定价公开($0.042/M)、OpenJev 生态提供自托管路线。Luna 的反制点是更简单的请求形状、今天就支持图片上下文,以及 OpenAI 第一方集成——代价是预览阶段的定价与校准透明度。
现在上生产应该选哪个?
在 Luna 还是 limited preview 的当下默认选 Jev:它正式可用、阈值经过校准、定价公开,而且迁移只是包一层适配函数。三类情况优先看 Luna:技术栈 all-in OpenAI、决策需要图片上下文、或单一厂商账单比价格透明更重要。无论选谁,先拿约 100 条自己业务的标注样本跑过两个闸门,再谈「只凭置信度自动执行」。
能不能两个都在同一条管道里用?
可以,而且模式是现成的:把两者都收口到代码里的同一个决策接口后面,再让置信度门控降级链负责路由——Jev 闸门常驻第一道,需要图片上下文的决策走 Luna,同一条 0.60–0.85 复核带接住任何一家的低置信输出。这样换厂商的成本是一个适配器,而不是一次重写。
这和 Jev vs Luna 基准页有什么区别?
本页对比的是产品与 API 契约:请求形状、原语、延迟、定价、校准、部署。Jev vs Luna 基准页对比的是实测任务成绩——505 样本第三方评测中 Jev 以约 $0.01 对 $0.06 的成本拿到 382/505,其中也包括 Luna 赢的那一局。两页一起读:契约决定你能建什么,基准决定建出来该期待什么。
预览版契约变了,我的代码怎么办?
OpenAI 承诺「扩大开放时披露更多细节」,所以要假设请求与响应字段都可能变——这正是 Luna 调用必须收口在一个适配函数里、应用层永远不直接看见它的原因。Jev 的 evaluate 端点自上线以来正式可用且未变。无论如何把手头那约 100 条评测集留着:换一个闸门重新验证是一个下午的事,不是一个季度。