跳到正文
openrouter blog·· 10 天前精選AI 評分63

Jev 與 Claude Opus 5 語意分類準確度與效能評測

Is Jev as Accurate as Frontier Models at Classification?

AI 導讀

OpenRouter 測試了專用決判模型 Jev 1.13 與前沿聊天模型 Claude Opus 5 在 Banking77 語意分類任務上的表現。

推薦理由

文章對比了專用決判模型與前沿聊天模型在分類任務上的表現,讀者可參考其評測資料與成本差異評估應用方案。

正文 · AI 翻譯

譯文尚未完整,完整內容請切換至原文。

Jev 是 TypeSafe 的 System One 決策模型。給它一個應用程式狀態物件和一個已型別化的問題,它會回傳一個已型別化的答案、一個置信度分數,以及每個可能選項的機率(即 Decisions API 的回應型別,而非聊天端點)。根據 TypeSafe 的說明,你可以用像 Jev 這樣的小型判斷模型取代前沿聊天模型,在分類任務(如本例)上同時大幅降低成本與延遲。你會犧牲多少準確度?

We ran 3,080 Banking77 customer support utterances through both Jev 1.13 and Claude Opus 5, the model OpenRouter users spent the most on for classification in the Task spend section of its rankings page as of 22 September 2026. Each utterance gets sorted into one of 77 banking intents, and both models got the same one-line description of each intent.

下方圖表將 Jev 與 Opus 置於同一畫面,顯示它們在準確度、中位延遲與每千次請求成本上的比較。

Three bar charts comparing Jev 1.13 and Claude Opus 5 on the Banking77 test split. Accuracy is 81.0% for Jev and 84.4% for Opus. Median latency is 175 ms for Jev and 2,266 ms for Opus. Cost per 1,000 requests is $0.11 for Jev and $2.42 for Opus. Title reads that Jev trails Claude Opus 5 by 3.3 points at 13x the speed and 1/22 the cost.

Jev 的準確度比 Opus 落後 3.3 分,但其中位延遲快 13 倍,且成本僅為 Opus 的 1/22。

Jev vs Claude Opus 5 一覽

以下是更詳細的結果。Macro-F1 對每個意圖給予相等權重。

Jev 1.13Claude Opus 5
Accuracy81.0% (79.6 to 82.3)84.4% (83.1 to 85.6)
Macro-F180.5%83.6%
Invalid responses0 of 3,0800 of 3,080
Latency p50175 ms2,266 ms
Latency p95270 ms3,004 ms
Latency p99353 ms3,835 ms
Billed cost, full run$0.34$7.44
Cost per 1,000 requests$0.11$2.42
Mean input tokens2,6053,750 (3,725 cached)
Mean output tokens82617

括號中的數字是 95% 的 bootstrap 信賴區間,我們使用 3,080 個範例計算。兩個模型皆未回傳格式錯誤的回應。每一次錯誤都是標籤錯誤,而非解析失敗。

How we ran it

Banking77 是 PolyAI 的一個以語句為單位的客戶支援意圖資料集,採用 CC BY 4.0 授權。這些語句短小,平均九個單字,且每句皆標註 77 個意圖之一。當我們在 Banking77 上測試 Jev 與 Opus 時,我們使用完整測試集,共 3,080 個範例,平均每個意圖 40 個。部分意圖相近,模型無法僅依靠少數關鍵字就完成判斷,需閱讀整句。請參考 card_arrival 與 card_delivery_estimate。

我們僅根據標籤名稱為每個意圖撰寫了一行標準,完全沒有參考測試資料,並將相同的 77 個標準列表提供給 Jev 和 Opus。為了評分 Jev,我們將每句話作為 Decisions API Choice 問題傳送,將 77 個標準作為選項。對於 Opus,我們構造了一個提示,系統訊息列出所有標準,然後使用者訊息僅包含該句話。Opus 以溫度 0 執行,關閉推理,並使用嚴格的 JSON schema 回應格式,列舉 77 個標籤。由於系統訊息在每次請求中相同,我們為其開啟了 提示快取。兩個模型均在同一臺機器上執行,我們一次傳送 8 個併發請求,並以客戶端觀測到的 OpenRouter 來測量往返延遲。

How accurate is Jev?

讓我們談談數字。於 Banking77 上,Jev 的準確度為 81.0%,Opus 為 84.4%。配對 bootstrap 給出 95% CI 為 2.3 至 4.4 分,Opus 以約三分之差領先,並非噪音。

另一方面,兩者之間有大量一致。它們在 89.3% 的語句上達成一致。當分歧時,Opus 單獨正確 175 筆,Jev 正確 72 筆。

另一方面,兩者仍遠低於在所有 10,003 個 Banking77 訓練範例上微調的編碼器所報告的低 90% 以上。這是依賴單行條件而非完整微調所付出的代價。

按類別來看,Opus 在 77 個意圖中領先 Jev 35 個,Jev 領先 15 個,平手 27 個。Opus 在 receiving_money 上的領先幅度最大,達 80.0% 相較於 Jev 的 52.5%。同時,Jev 在 compromised_card 上表現優異,達 95.0%,而 Opus 只有 70.0%。Opus 傾向於將被盜卡資訊誤判為未識別付款。

Jev 的速度與成本如何?

Jev 的中位數往返時間僅 175 毫秒,95 分位數為 270 毫秒。相對地,Opus 在關閉推理模式下的中位數往返時間為 2,266 毫秒,95 分位數為 3,004 毫秒。這表示最慢的 Jev 呼叫(約 1.6 秒)比最快的 Opus 呼叫(約 1.9 秒)還快。

Jev 的總成本為 0.34 美元,亦即每千次請求 0.11 美元。Opus 的成本為 7.44 美元,亦即每千次請求 2.42 美元。此 Opus 數字已將 3,700 代幣的系統提示快取,故所有請求皆以快取讀取速率計費,而非列表速率。若不使用快取,Opus 的列表速率約為每千次請求 19 美元。若您使用前沿模型的標籤分類法,請將該系統提示快取。

根據 Jev 的信心值進行路由

Jev 會給出一個信心分數,但並非經過校準的機率。此處它在中等範圍內高估了自身準確度。儘管如此,排名仍相當不錯。對於 58% 的語句,其信心值 ≥0.99,準確率為 96.3%。對於低於 0.5 的 3.5%,準確率為 29.6%。

這個排序就是你需要的一切,以進行cascade。在給定閾值以上接受 Jev,將其餘送至 Opus。以下列出了各閾值下的準確率與每千次成本,範圍從僅使用 Jev 到僅使用 Opus。

閾值由 Jev 處理準確率每千次成本
僅 Jev100%81.0%$0.11
0.9957.8%84.3%$1.13
0.9568.1%84.2%$0.88
0.9075.9%84.0%$0.69
0.8082.7%83.6%$0.53
0.7087.3%83.2%$0.42
僅 Opus0%84.4%$2.42

在閾值 0.90 時,76% 的流量直接使用 Jev,避免了 Opus。相對於僅使用 Opus,您的準確率最多下降 0.4 分,成本則減少 3.5 倍。Opus 正確處理了 Jev 遺漏的 175 條語句。這些案例的中位信心值為 0.67,其中 85% 的信心值低於 0.90。因此,Jev 通常能知道何時僅在猜測。

注意事項

讓我們具體討論我們進行的實際測試:一次資料集、一次領域、一次提示設計,以及一次在某個下午的十五分鐘視窗。

若 Opus 在此處具有記憶優勢,其分數可能會因為 Banking77 於 2020 年發布而略顯高估。但僅憑此一次執行無法判斷。

另一個重點是,兩個模型在同一個類別上失誤,因為條件僅根據標籤名稱編寫,未參考範例訊息。在此情況下,get_physical_card 標籤涵蓋了是否單獨寄送 PIN 的查詢。僅從名稱推斷幾乎不可能。兩個模型在此類別上皆得 0 分(共 40 分),將大多數訊息路由至 change_pin。排除該類別後,Jev 的準確率為 82.1%,Opus 為 85.5%。然而,發布團隊會使用保留的訓練資料對條件進行迭代,而非測試集,兩個模型將會提升。

我們展示的 cascade 表格使用了與準確率測量相同的 3,080 個範例,因此這是一個上限。請根據自己的流量選擇閾值。

最後,這些測試是在 Opus 關閉推理模式且開啟結構化輸出時進行的。Opus 未在「開啟推理」模式、其他架構或少量示例下進行測試。

這意味著

如果在準確率上比每千次請求 $2.42 更重要,請選擇 Opus。若您處理高量、低延遲或需要備援槽位,單獨使用 Jev 的準確率僅落後 Opus 3.3 點。更好的是,在 Jev 的信心分數上執行級聯,我們僅以不到 30% 的成本就能達到距離 Opus 0.4 點的準確率。

The Decisions API reference walks through the basic request shape. The Jev cookbook walks you through a working Choice question in TypeScript. The Jev-verified cascade cookbook has the escalation pattern in code, with a cheap model drafting, Jev checking the draft, and a frontier model handling only what fails the check. The Banking77 test split is a 3,080-line CSV in the PolyAI repository.

To run a labeling job like this one over your own backlog, the Classify and Tag Text at Scale with Jev cookbook covers batching within rate limits, picking a threshold per label from a labeled sample, and computing cost per 1,000 items. If Jev is new to you, start with What Is Jev?, then Jev vs LLM for when a decision model replaces a generative call. The Jev documentation hub lists every Jev guide and cookbook on OpenRouter.

Frequently Asked Questions

How accurate is Jev compared to Claude Opus 5 on intent classification?

On the Banking77 test split consisting of 3,080 utterances across 77 intents run on 22 September 2026, Claude Opus 5 achieved 84.4% accuracy and 83.6% macro-F1, while Jev 1.13 achieved 81.0% accuracy and 80.5% macro-F1. The paired gap was 3.3 percentage points on accuracy with a 95% bootstrap interval from 2.3 to 4.4 percentage points. The models agreed on 89.3% of utterances.

How much faster is Jev than a frontier model for classification?

On median and p95 client-observed round trip latency, Jev achieved median 175 ms and p95 270 ms, while Claude Opus 5 with reasoning disabled achieved median 2,266 ms and p95 3,004 ms, about 13 times higher at the median. These numbers are measured from the same client under 8 concurrent requests. They include network time and any queueing in the provider’s system before the actual inference.

How much does it cost to classify text with Jev on OpenRouter?

All 3,080 of the Banking77 requests were billed $0.34 total on typesafe/jev-1.13, which is $0.11 per thousand requests or about $0.0001 per request. Using prompt caching on the 77-label system prompt, the same requests billed $7.44 total on anthropic/claude-opus-5, or $2.42 per thousand. Without caching, Opus would have billed about $19 per thousand at list price.

Can I use Jev’s confidence score to decide when to fall back to a bigger model?

Yes, and in our data it works as a ranking signal. We measured Jev as 96.3% accurate on the 58% of utterances with confidence at least 0.99, and 29.6% accurate on the 3.5% of utterances with confidence under 0.5. Routing everything under 0.9 confidence to Claude Opus 5 recovered 84.0% accuracy, within 0.4 points of Opus alone, at $0.69 per thousand requests. Because the score is not calibrated as a probability, pick the threshold that works for your own data.

來源:openrouter blog · openrouter.ai