Jev Recipe / 厂商对比
Jev vs Clef:Cloudflare 决策模型正面比较
Cloudflare 用开源的 Clef 回应了 Jev。带出处的记分牌对比:BFCL 与 When2Call 的分裂、38.8ms 托管延迟、Apache 2.0 开源权重,以及「fully Jev-API compatible」这个声明到底承诺了什么。
选定决策模型之前的 6 项检查
{
"routing": {
"type": "choice",
"instructions": "Which queue should this ticket go to?",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Product malfunction or bug",
"sales": "Buying or upgrade question",
"abuse": "Safety or abuse report"
}
},
"needs_safety_escalation": {
"type": "noul",
"instructions": "Does this ticket require a safety escalation regardless of queue?"
},
"urgency": {
"type": "score",
"instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)"
}
}决策模型记分牌:TypeSafe Jev vs Cloudflare Clef(发布周,逐格标注来源)
代码对比:Jev 类型化调用 vs Clef 本地运行与 A/B 切换
import requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_ROUTE_CONFIDENCE = 0.85
QUESTIONS = {
"routing": {
"type": "choice",
"instructions": "Which queue should this ticket go to?",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Product malfunction or bug",
"sales": "Buying or upgrade question",
"abuse": "Safety or abuse report",
},
},
"needs_safety_escalation": {
"type": "noul",
"instructions": "Does this ticket require a safety escalation regardless of queue?",
},
"urgency": {
"type": "score",
"instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
},
}
resp = requests.post(
JEV_ENDPOINT,
json={"state": {"ticket": "..."}, "questions": QUESTIONS},
timeout=5,
)
resp.raise_for_status()
data = resp.json()
route = data["routing"] # 类型化答案 + 校准置信度,零生成
if (
route["confidence"] >= AUTO_ROUTE_CONFIDENCE
and not data["needs_safety_escalation"]["answer"]
):
lane = f"auto:{route['answer']}" # 单次前向,仅输入计费
else:
lane = "review" # 0.60~0.85 带宽或安全标记import json
import requests
# 来自本站 Clef-Flash 装机实测的已验证本地路径:
# bartowski Q4_K_M GGUF、Ollama 0.35、8GB 笔记本、2048 上下文。
# 机制注记:这条路径是「生成 JSON」(21–30 个 token)——官方
# joint schema head 未被加载(427 个 backbone 张量、零
# schema head 匹配),因此没有原生概率。
OLLAMA_CHAT = "http://localhost:11434/api/chat"
SCHEMA = {
"type": "object",
"properties": {
"urgent": {"type": "boolean"},
"team": {"enum": ["billing", "technical", "sales"]},
"severity": {"enum": ["minor", "major", "critical"]},
},
"required": ["urgent", "team", "severity"],
}
resp = requests.post(
OLLAMA_CHAT,
json={
"model": "hf.co/bartowski/Cloudflare_clef-flash-GGUF:Q4_K_M",
"messages": [{
"role": "user",
"content": "Ticket: Checkout has been failing for every "
"customer for the last hour. Reply with urgent, team, "
"severity as JSON.",
}],
"format": SCHEMA, # 运行时层约束修的是「格式」,
"options": {"temperature": 0, "num_ctx": 2048}, # 不是「正确性」
},
timeout=60,
)
reply = json.loads(resp.json()["message"]["content"])
# 热请求约 1.4~2 秒。我们的测试里,同一句宕机提示词零温度复跑
# 把 team 从 technical 翻成了 billing——路由评测要有标注样本集,
# 不能靠一张成功截图。const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_ROUTE_CONFIDENCE = 0.85;
const QUESTIONS = {
routing: {
type: "choice",
instructions: "Which queue should this ticket go to?",
criteria: {
billing: "Payment, invoice or refund issue",
technical: "Product malfunction or bug",
sales: "Buying or upgrade question",
abuse: "Safety or abuse report",
},
},
needs_safety_escalation: {
type: "noul",
instructions:
"Does this ticket require a safety escalation regardless of queue?",
},
urgency: {
type: "score",
instructions:
"Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
},
} as const;
const resp = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state: { ticket: "..." }, questions: QUESTIONS }),
});
const data = await resp.json();
const lane =
data.routing.confidence >= AUTO_ROUTE_CONFIDENCE &&
!data.needs_safety_escalation.answer
? `auto:${data.routing.answer}`
: "review";// Clef 官方自称 "fully Jev-API compatible"——如果对你的调用成立,
// 换厂商就是换这个常量。保留意见:这是 Cloudflare 的自述,本站
// 未逐端点验证;而且它描述的是托管模型(本地量化路径生成 JSON、
// 无原生决策头)。把两个后端都藏在同一个接口后面,先在约 100 条
// 标注样本上验证切换,再让自动化依赖置信度数字。
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const DECISION_ENDPOINT =
process.env.DECISION_BACKEND === "clef"
? process.env.CLEF_ENDPOINT! // 你的 Workers AI / 自托管 Clef 端点
: JEV_ENDPOINT;
const QUESTIONS = {
routing: {
type: "choice",
instructions: "Which queue should this ticket go to?",
criteria: {
billing: "Payment, invoice or refund issue",
technical: "Product malfunction or bug",
sales: "Buying or upgrade question",
abuse: "Safety or abuse report",
},
},
} as const;
export async function routeTicket(ticket: string) {
const resp = await fetch(DECISION_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state: { ticket }, questions: QUESTIONS }),
});
if (!resp.ok) throw new Error("decision call failed");
const data = await resp.json();
// 各后端的置信度语义可能不同:自动化闸门建立在「按厂商逐个
// 验证过的校准」上,而不是「存在一个数字」上。
return data.routing as { answer: string; confidence: number };
}Jev vs Clef 常见问题
Clef 是什么?和 Jev 是什么关系?
Cloudflare 于 2026-10-01 发布的开源权重决策模型家族:Clef(27B)与 Clef-Flash(9B),从冻结的 Qwen3 底座加 rank-256 LoRA 微调而来,以 Apache 2.0 许可上架 Hugging Face,并由 Workers AI 托管。与 TypeSafe 的 Jev 一样它是非生成型的:在单次 prefill 并行前向里给预定义选项打分,并官方自称「fully Jev-API compatible」。发布当天在 Hacker News 拿下 625 分——本站监测史以来最大单帖——「Jev 兼容」从此从社区复现升级为一线大厂产品动作。
「Jev-API compatible」到底是什么意思?
按 Cloudflare 公告口径,Clef 端点应当接受相同的请求形状——上下文加 questions 加预定义答案集——并返回答案加置信度,于是把 Jev 换成 Clef 是改一个端点而不是重写。两个保留意见:这是厂商自述,本站没有逐端点验证;而且它描述的是托管模型——本地量化路径跑的是另一种机制(生成 JSON、无原生决策头),形状兼容不等于置信度语义兼容。
该选谁?
按场景选,不按记分牌选。Clef 优先:你住在 Cloudflare 生态里、今天就要图像输入、想要可自托管的开源权重,或在意 BFCL/BANKING77 领先(格式精确的函数调用、银行精度)与 ~38.8ms 托管延迟。Jev 优先:你要靠置信度阈值做自动化、需要 RLCD 校准概率、输出免费的仅输入计费、多问题一次调用,或 When2Call 80.97 对 65.58 的画像(带弃权的工具选择)。无论哪边:自动化之前先用约 100 条自己的标注样本过一遍闸门——发布周厂商数字回答的是选型侦察问题,不是部署问题。
Clef 能本地跑吗?
能,但要先问机制。bartowski 的 Clef-Flash Q4_K_M GGUF 在 8GB RTX 4060 笔记本的 Ollama 0.35 里热请求约 1.4–2 秒一次,我们的测试里它也接受了图像输入。但受检量化包只有 427 个 backbone 张量、零 schema head 匹配:Ollama 的答案是生成的 JSON(21–30 个 token),不是官方 joint schema head 的原生概率。要原生决策接口,参考路径是 cloudflare/clef-webcam 仓库(Apple Silicon、32GB+ 内存)。经验法则:查运行时——模型名字不会告诉你跑的是哪种机制。
Clef 免费吗?许可证是什么?
权重以 Apache 2.0 上架 Hugging Face,下载与自托管免费——成本是你的 GPU。Cloudflare 在 Workers AI 上的托管是付费产品,其按 token 计价在撰稿时可核实的信源中未公布(2026-10-05)——如实记为「未公布」而不是免费。Jev 的对照点是公开的每百万输入 token $0.042、输出免费。
基准分数分裂该怎么读?
先问每个基准考什么。BFCL(98.76 对 95.75)与 API-Bank(93.11 对 88.19)奖励格式精确的函数调用——Clef-Flash 领先。BANKING77 macro-F1(94.20 对 79.74)奖励细粒度银行意图精度——Clef 优势最大的一格。When2Call(80.97 对 65.58)奖励选对工具或弃权——Jev 领先最多的一格,工作流层分数也大多跟随(发票处理 61.8 对 57.1、agent traces 71.6 对 69.8、客服则是 77 对 76 的掷硬币)。以上全部是 Cloudflare 公开评测;Jev 的独立数字在本站基准页。分裂本身就是结论:这两个模型擅长不同的工作。