跳到正文
openrouter blog·· 13 天前精選AI 評分77

Jev 與 LLM 搭配指南:何時使用決策模型而非生成文本

Jev vs LLM: When to Use a Decision Model Instead of Generating Text

AI 導讀

本文探討如何結合 TypeSafe 的決策模型 Jev 1.13 與大型語言模型建構客服自動化管線,透過分流與驗證機制兼顧成本與準確率。

  • 效能與成本:在支援票證分類測試中,Jev 1.13 的成本僅為 $0.0248 每千筆,延遲約 194 ms,準確率與 GPT Luna 相當。
正文 · AI 翻譯

假設你有一個產品每個月收到 40,000 張支援工單,每一張都需要進行分類,提供三項資訊:問題內容、是否需要升級,以及回覆內容。

一個前沿的 LLM 搭配 JSON 架構即可完成此工作。我們的基準測試顯示,其成本為每 1,000 張工單 2.88 美元,平均延遲兩秒,約每月 115 美元。較小的 LLM 成本則為每 1,000 張 0.09 美元,平均延遲約一秒。

另一種思考方式是詢問決策模型能做什麼。Jev 是 TypeSafe 的一個模型,在同一次測試中,其分類成本為每 1,000 張工單 2.5 分,平均延遲 194 毫秒,約每月一美元。且它會回傳機率,無需額外解析。

但它無法撰寫回覆。這部分仍需使用 LLM。因此,我們在 OpenRouter 上針對 100 個支援案例測試了 Jev、GPT Luna 與 Claude Opus,並構建了兩種結合兩者的模式。

Jev 是什麼,以及每個模型實際回傳什麼

要求語言模型做程式碼決策時,它會回傳文字。你可以要求 JSON,但模型本身是為寫作設計的。你的程式需要解析輸出,確保模型是回答而非說明。

Jev 則回傳一個有型別的答案。你傳入狀態與由三種原語構成的問題:choice 從你定義的集合中選擇一項,noul 回傳「是」的機率,score 將內容放置於你描述的層級。選項與分數答案可包含你自訂鍵的機率,而 noul 為單一數值。以下為兩個關於一張帳單工單的回覆。

{
  "answers": {
    "intent": {
      "type": "choice",
      "choice": "billing_dispute",
      "confidence": 1,
      "probabilities": { "billing_dispute": 1, "order_status": 0, "other": 0 }
    },
    "escalate": { "type": "noul", "noul": 0.04 }
  },
  "model": "typesafe/jev-1.13-20260917",
  "provider": "TypeSafe",
  "usage": { "cost": 0.000016002, "inputTokens": 381, "outputTokens": 62 }
}

意圖是你其中一個鍵值,升級則為你自行設定閾值所比對的數字。

TypeSafe 將 Jev 稱為 System One model。這表示它做出快速、直觀的判斷,而非慢速、深思熟慮的分析。它僅接受文字輸入,不能處理影像、音訊或 PDF。它會逐字讀取標準。更重要的是,當其檢查的狀態包含無關細節時,準確性會下降,並且在算術、計數與日期比較方面不可靠。

將 Jev 與生成式 LLM 並排比較

Jev 1.13傳統 LLM(GPT Luna、Claude Opus)
輸出有型別的選擇、是/否機率或分數,並包含每項選項的機率自由文本(可選限制為 JSON)
最佳適用場景分類、路由、驗證、排名、守護欄檢查、受限提取寫作、解釋、摘要、重寫、程式碼,以及任何輸出空間事先未知的情況
輸入僅文字(字串、JSON、陣列)。每次請求 64k 令牌,狀態加最長問題共 32k 令牌文字,且許多模型支援影像、音訊、PDF
價格(2026-09-19 觀測)$0.042 每百萬輸入令牌,輸出免費GPT Luna $0.20 入口 / $1.20 出口 每百萬。Claude Opus $5 入口 / $25 出口 每百萬
60 張工單分類的中位延遲194 毫秒1,106 毫秒(Luna),1,957 毫秒(Opus)
每 1,000 張已分類工單的成本$0.0248$0.0921(Luna),$2.88(Opus)
OpenRouter 端點POST /api/alpha/decisionsPOST /api/v1/chat/completions

價格來自 OpenRouter 模型後設資料(執行日期)。請在預算前檢視模型頁面。

我們測量的內容

模型為 Jev 1.13、GPT Luna(解析為 GPT 5.6 Luna)和 Claude Opus(解析為 Claude Opus 5)。LLM 在系統提示中獲得與 Jev 相同的定義和升級規則,溫度 0,並請求返回純 JSON 物件。每個示例僅一次請求。資料集規模較小,分別為 60 和 40 個示例,均在同一天使用同一組提示。這裡展示的是權衡關係,而非排行榜。

任務 A:將 60 筆客服工單分為五個意圖並加上升級標誌

六十筆客服工單涵蓋五個意圖:訂單狀態、退貨或退款、帳單爭議、產品問題、帳戶存取,每項都有一句定義。升級範圍包含法律威脅、退費、疑似詐騙、安全風險,以及公開威脅。

模型意圖準確率升級準確率p50 延遲p95 延遲總成本 (60)每 1,000
Jev 1.1359/60 (98.3%)60/60 (100%)0.194s0.633s$0.001489$0.0248
GPT Luna59/60 (98.3%)60/60 (100%)1.106s2.395s$0.005527$0.0921
Claude Opus60/60 (100%)59/60 (98.3%)1.957s2.594s$0.17281$2.8802

準確率相差不大。Jev 送出的輸入 token 更多,因為每個請求都帶全套條件,仍然只比 GPT Luna 費用低約四分之一、Claude Opus 低於百分之一。只有一次意圖錯誤,關於退貨商品退款分拆的問題。Jev 以 0.56 的置信度將其歸類為退貨或退款。它是唯一低於 0.8 的工單,因此即使設 0.8 門檻,也會被送交人工處理。

任務 B:篩選 40 條訊息以檢測 prompt 注入

二十二條普通客服訊息與十八條注入嘗試(忽略指令、偽 admin 標記、角色扮演框架、引用指令)。我們給 Jev 一個是/否問題:這是否試圖改變助手行為?LLM 取得相同定義並回傳布林值。

模型準確率p50 延遲p95 延遲總成本 (40)每 1,000
Jev 1.1340/40 (100%)0.194s0.688s$0.000646$0.0161
GPT Luna39/40 (97.5%)0.805s1.491s$0.001828$0.0457
Claude Opus40/40 (100%)2.099s4.377s$0.063555$1.5889

Jev 將兩組分離得很乾淨。注入訊息得分 0.86 至 0.99,普通訊息 0.01 至 0.20。Luna 唯一錯誤是角色扮演框架。

當答案是 N 個選項之一時使用 Jev

任何以 switch 陳述式結尾的都是 Jev 問題。意圖、優先順序、語言、情緒、排程、此舉是否違反政策、此主張是否得到支援。每種情況答案都已預先知道。你只需要各可能答案的機率。以下為 benchmark 的 triage 呼叫,使用 OpenRouter TypeScript SDK。

// triage.ts
import { OpenRouter } from '@openrouter/sdk';

const openrouter = new OpenRouter({
  apiKey: process.env.OPENROUTER_API_KEY, // server-side only
});

const INTENTS = {
  order_status:
    'The customer asks where an order is, when it will ship or arrive, wants to change or cancel an order before delivery, or reports a package missing or partially delivered.',
  return_refund:
    'The customer wants to return, exchange, or replace an item they received, or asks about the status or rules of a return they already started.',
  billing_dispute:
    'The customer says a charge, invoice, tax, discount, or refund amount is wrong, duplicated, unexpected, or unauthorized.',
  product_question:
    "The customer asks about a product's features, compatibility, sizing, materials, stock, warranty, or safety before or after buying, without asking to return it.",
  account_access:
    'The customer cannot log in, needs to change login or account details, or asks to merge, delete, secure, or share an account.',
} as const;

export type Intent = keyof typeof INTENTS;

const ESCALATE =
  'Does the ticket describe any of the following: a threat of legal action, a regulator complaint, or a chargeback; suspected fraud or an account takeover; a safety hazard such as fire, smoke, or injury; or a customer who says this is a repeated failure and threatens to publicize it?';

export type Triage = {
  intent: Intent;
  confidence: number;
  probabilities: Record<string, number>;
  escalateProbability: number;
  costUsd: number;
};

function isIntent(value: string): value is Intent {
  return Object.hasOwn(INTENTS, value);
}

export async function triage(ticket: string): Promise<Triage> {
  const result = await openrouter.alpha.decisions.create({
    decisionsRequest: {
      model: 'typesafe/jev-1.13',
      state: { ticket },
      questions: {
        intent: {
          type: 'choice',
          instructions: 'What is the primary intent of the ticket?',
          criteria: INTENTS,
        },
        escalate: { type: 'noul', instructions: ESCALATE },
      },
    },
  });

  const intent = result.answers.intent;
  const escalate = result.answers.escalate;
  if (intent?.type !== 'choice' || escalate?.type !== 'noul') {
    throw new Error('Unexpected answer types');
  }
  if (!isIntent(intent.choice)) {
    throw new Error(`Unknown intent ${intent.choice}`);
  }
  return {
    intent: intent.choice,
    confidence: intent.confidence ?? 0,
    probabilities: intent.probabilities ?? {},
    escalateProbability: escalate.noul,
    costUsd: requireCost(result.usage.cost),
  };
}

function requireCost(cost: number | undefined): number {
  if (cost === undefined) {
    throw new Error('Response did not include usage.cost');
  }
  return cost;
}

該 triage 程式碼中需注意的三點:

每題只提出一個小判斷,並在程式碼中合併答案。 在此,意圖與升級是同一請求中的獨立問題。

將條件寫成規格式。 由於 Jev 逐字執行,若案例落入錯誤分類,先修正條件文字。

僅傳遞每題所需的狀態。 不相關細節會降低答案準確率

當答案為散文時使用 LLM

現在當我們需要實際產生回覆時,就使用 LLM。

// draft-reply.ts
import { OpenRouter } from '@openrouter/sdk';

const openrouter = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY });

const REPLY_MODEL = '~openai/gpt-luna-latest';

const REPLY_SYSTEM =
  'You write short, warm replies for the Northwind support team. Use only the facts in the user message. Do not promise refunds, credits, or dates that are not in the facts. Three sentences maximum.';

type Draft = { text: string; model: string; costUsd: number };

export async function draftReply(ticket: string, facts: string): Promise<Draft> {
  const result = await openrouter.chat.send({
    chatRequest: {
      model: REPLY_MODEL,
      messages: [
        { role: 'system', content: REPLY_SYSTEM },
        { role: 'user', content: `Ticket: ${ticket}\n\nFacts: ${facts}` },
      ],
    },
  });
  // chat.send can return a stream when streaming is requested. It isn't here, so reject that case.
  if (result instanceof ReadableStream) {
    throw new Error('Expected a non-streaming response');
  }
  const text = result.choices[0]?.message.content;
  if (typeof text !== 'string') {
    throw new Error('Expected text content');
  }
  const cost = result.usage?.cost;
  if (typeof cost !== 'number') {
    throw new Error('Response did not include usage.cost');
  }
  return { text, model: result.model, costUsd: cost };
}

依寫作任務選擇模型。若需短、快速、低成本回覆,GPT Luna 便能交付,本次執行約每回覆 $0.0001。若寫作需要實際推理、長上下文或程式碼,則升級為前沿模型。當遇到 Jev 的限制(如輸入包含影像、音訊或 PDF,或輸出為檔案、差異或計畫)時,LLM 亦是答案。

使用 Jev 路由,程式碼計算,LLM 撰寫

從宏觀來看,第一種模式很簡單。Jev 進行分流。程式碼檢查其確定度。僅在交付物為散文時,程式碼才呼叫 LLM。

處理器將任何升級門檻 0.5 或以上的訊息送至人工。亦將意圖置信度低於 0.8 的工單送交人工。否則,它直接從訂單系統回答訂單狀態,無需模型,且僅對需散文的意圖請求 LLM 產生草稿。

// handle.ts
import { triage, type Intent, type Triage } from './triage';
import { draftReply } from './draft-reply';

// These thresholds are application policy. Tune them on your own labeled tickets.
const ROUTE_CONFIDENCE = 0.8;
const ESCALATE_THRESHOLD = 0.5;

type ReplyIntent = Exclude<Intent, 'order_status'>;

type Route =
  | { kind: 'human'; reason: string }
  | { kind: 'deterministic'; intent: 'order_status' }
  | { kind: 'reply'; intent: ReplyIntent };

function chooseRoute(t: Triage): Route {
  if (t.escalateProbability >= ESCALATE_THRESHOLD) {
    return { kind: 'human', reason: `escalation probability ${t.escalateProbability}` };
  }
  if (t.confidence < ROUTE_CONFIDENCE) {
    return { kind: 'human', reason: `intent confidence ${t.confidence}` };
  }
  if (t.intent === 'order_status') {
    return { kind: 'deterministic', intent: t.intent };
  }
  return { kind: 'reply', intent: t.intent };
}

// Stand-in for your order system. Exact data never goes through a model.
function lookupOrder(ticket: string): string {
  const id = ticket.match(/\b\d{5}\b/)?.[0];
  if (id === undefined) {
    return 'Please reply with your five-digit order number and we will check the shipment.';
  }
  return `Order ${id} shipped 2026-09-17 via UPS, tracking 1Z999AA10123456784, estimated delivery 2026-09-21.`;
}

export const FACTS: Record<ReplyIntent, string> = {
  return_refund:
    'Returns are accepted within 30 days of delivery for regular items. Exchanges for another size follow the same 30-day window. Final sale items cannot be returned or exchanged. Prepaid labels are emailed within one business day.',
  billing_dispute: 'A billing specialist will review the charge within one business day.',
  product_question: 'Product specifications are listed on each product page.',
  account_access: 'Password resets are available from the sign-in page.',
};

export type Handled = {
  triage: Triage;
  route: Route;
  text: string;
  costUsd: number; // Jev call plus the LLM call, if one happened
};

// Routes a ticket. A 'reply' route carries an unverified LLM draft; send.ts checks it before anything goes out.
export async function handle(ticket: string): Promise<Handled> {
  const t = await triage(ticket);
  const route = chooseRoute(t);
  switch (route.kind) {
    case 'human':
      return { triage: t, route, text: `queued for a person: ${route.reason}`, costUsd: t.costUsd };
    case 'deterministic':
      return { triage: t, route, text: lookupOrder(ticket), costUsd: t.costUsd };
    case 'reply': {
      const reply = await draftReply(ticket, FACTS[route.intent]);
      return { triage: t, route, text: reply.text, costUsd: t.costUsd + reply.costUsd };
    }
    default:
      return route satisfies never;
  }
}

簡易執行器會在每個結果旁印出 Jev 訊號。

// run.ts
import { handle } from './handle';

const TICKETS = [
  "Hi, I ordered a pair of running shoes on the 3rd and the tracking page hasn't updated in five days. Where is my package? Order 84721.",
  'The jacket I got is too small. Can I swap it for a large? I got it last Tuesday.',
  "There's an unauthorized $450 charge from your company on my card. I've already called my bank to dispute it and I'm reporting this as fraud.",
  "I paid with a gift card and a credit card, but the refund only went to the credit card. Where's the gift card balance?",
];

for (const ticket of TICKETS) {
  const r = await handle(ticket);
  const { intent, confidence, escalateProbability } = r.triage;
  console.log(`"${ticket}"`);
  console.log(`  intent=${intent} confidence=${confidence} escalate=${escalateProbability} route=${r.route.kind} cost=$${r.costUsd.toFixed(6)}`);
  console.log(`  ${r.text}\n`);
}

以下為四張範例工單的輸出。

"Hi, I ordered a pair of running shoes on the 3rd and the tracking page hasn't updated in five days. Where is my package? Order 84721."
  intent=order_status confidence=1 escalate=0.02 route=deterministic cost=$0.000026
  Order 84721 shipped 2026-09-17 via UPS, tracking 1Z999AA10123456784, estimated delivery 2026-09-21.

"The jacket I got is too small. Can I swap it for a large? I got it last Tuesday."
  intent=return_refund confidence=1 escalate=0.01 route=reply cost=$0.000118
  Yes, you can exchange the jacket for a large if it’s a regular item and within 30 days of delivery. If it’s a final sale item, it can’t be exchanged; for eligible exchanges, a prepaid return label is emailed within one business day.

"There's an unauthorized $450 charge from your company on my card. I've already called my bank to dispute it and I'm reporting this as fraud."
  intent=billing_dispute confidence=1 escalate=0.96 route=human cost=$0.000025
  queued for a person: escalation probability 0.96

"I paid with a gift card and a credit card, but the refund only went to the credit card. Where's the gift card balance?"
  intent=return_refund confidence=0.54 escalate=0.02 route=human cost=$0.000025
  queued for a person: intent confidence 0.54

四個案例中有一個觸發了 LLM。訂單狀態查詢精確且免費。詐騙報告從未到達能夠保證任何結果的模型。由於分割退款票的信心(Jev 在 Task A 錯誤歸檔的那一張)低於 0.8,該票被送到人工處理。最後,外套回覆仍然是一份未檢查的草稿。檢查它即為第二種模式。

在傳送前,讓 Jev 驗證 LLM 的輸出

第二種模式流程相反。LLM 先起草,然後 Jev 在傳送前根據政策檢查草稿。這樣可以捕捉到違反政策的友好 AI 回覆。

// verify-draft.ts
import { OpenRouter } from '@openrouter/sdk';

const openrouter = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY });

export const POLICY =
  'Returns are accepted within 30 days of delivery for regular items. Final sale items cannot be returned. Refunds go back to the original payment method. Prepaid return labels are emailed within one business day of approval.';

type Label = 'supported' | 'unsupported' | 'declined';

export type Verdict = {
  label: Label;
  confidence: number;
  probabilities: Record<string, number>;
  costUsd: number;
};

function isLabel(value: string): value is Label {
  return value === 'supported' || value === 'unsupported' || value === 'declined';
}

export async function verifyDraft(policy: string, question: string, draft: string): Promise<Verdict> {
  const result = await openrouter.alpha.decisions.create({
    decisionsRequest: {
      model: 'typesafe/jev-1.13',
      state: { policy, customer_question: question, draft_reply: draft },
      questions: {
        support: {
          type: 'choice',
          instructions: 'How does draft_reply relate to policy and customer_question?',
          criteria: {
            supported:
              'The draft answers customer_question, and every fact, number, timeframe, and promise in it appears in policy.',
            unsupported:
              'The draft states a fact, number, timeframe, or promise that policy does not contain or contradicts, or it answers a different question.',
            declined:
              'The draft says policy does not cover the question and adds no facts of its own beyond what policy states.',
          },
        },
      },
    },
  });

  const support = result.answers.support;
  if (support?.type !== 'choice' || !isLabel(support.choice)) {
    throw new Error('Unexpected answer');
  }
  const cost = result.usage.cost;
  if (cost === undefined) {
    throw new Error('Response did not include usage.cost');
  }
  return {
    label: support.choice,
    confidence: support.confidence ?? 0,
    probabilities: support.probabilities ?? {},
    costUsd: cost,
  };
}

Jev 取得政策、客戶問題和草稿,接著回答一個單選題,判斷該草稿是否被支援、未支援或被拒絕。我們將四份草稿送入測試。第一份來自 GPT Luna,其他三份則為手寫稿,每份都針對特定分支撰寫。

Q: "How long until my refund shows up?"
Draft (GPT Luna): "Refunds are sent back to the original payment method, but we don't have a specific timeline for when they will appear. If your return is approved, a prepaid return label will be emailed within one business day."
  supported  confidence=0.09  { supported: 0.39, unsupported: 0.32, declined: 0.29 }

Q: "How long until my refund shows up?"
Draft (fabricated): "Refunds are processed within 5 to 7 business days after we receive the item."
  unsupported  confidence=1.00  { unsupported: 1, supported: 0, declined: 0 }

Q: "Where does my refund go?"
Draft (grounded): "Refunds go back to the original payment method, so it will return to however you paid for the order."
  supported  confidence=1.00  { supported: 1, unsupported: 0, declined: 0 }

Q: "How long until my refund shows up?"
Draft (grounded, wrong question): "Refunds go back to the original payment method, so it will return to however you paid for the order."
  unsupported  confidence=0.48  { unsupported: 0.65, supported: 0.14, declined: 0.21 }

虛構的時間線回覆被判定為未支援,分數 1.00;若該回覆已傳送,客戶在第八天會回信詢問退款位置。

Luna 的草稿才是有趣的,因為它把所有事實都與政策相符。但它一半拒絕、一半回答問題,並加入了客戶未詢問的細節。Jev 將機率分成三份,0.8 的閾值會將其送給客服代表快速檢視。

最後兩行使用相同的基礎草稿,但問了兩個不同的問題。注意,當草稿回答客戶所問的問題時,得分 1.00,即被支援;若回答不同的問題,則被判定為未支援。這正是驗證器應具備的功能,也說明瞭政策文本應使用與草稿相同的措辭。

我們最初嘗試了兩個是非題,結果混亂,因為正確說明政策不涵蓋此情況的草稿既是基礎又不是答案。相互排斥的結果促使我們改為單選題。

將所有流程連線起來

另一個檔案將路由優先流程連結到驗證器,確保 LLM 所寫的內容在傳送前都有被支援的判斷。

// send.ts
import { FACTS, handle, type Handled } from './handle';
import { verifyDraft, type Verdict } from './verify-draft';

const SEND_CONFIDENCE = 0.8;

type Outcome = Handled & { disposition: 'sent' | 'human_review'; verdict?: Verdict };

export async function process(ticket: string): Promise<Outcome> {
  const handled = await handle(ticket);
  switch (handled.route.kind) {
    case 'human':
      return { ...handled, disposition: 'human_review' };
    case 'deterministic':
      return { ...handled, disposition: 'sent' };
    case 'reply': {
      // Verify against the same facts the draft was written from.
      const verdict = await verifyDraft(FACTS[handled.route.intent], ticket, handled.text);
      const ok = verdict.label === 'supported' && verdict.confidence >= SEND_CONFIDENCE;
      return { ...handled, verdict, costUsd: handled.costUsd + verdict.costUsd, disposition: ok ? 'sent' : 'human_review' };
    }
    default:
      return handled.route satisfies never;
  }
}

如果驗證器退回一則本身看起來沒問題的回覆,請根據事實重新審視。事實往往忽略了客戶所詢問的內容。請檢視完整的請求流程。

ticket text
  |
  v
[Jev]  intent (choice) + escalate (noul) .......... 1 request, ~200 ms, ~$0.000025
  |
  v
[code] escalate >= 0.5 or confidence < 0.8 ?  -->  human queue
  |
  v
[code] intent == order_status ?  -->  database lookup, exact reply, no model
  |
  v
[LLM]  draft reply from ticket + facts ............ ~1 s, ~$0.0001 (GPT Luna)
  |
  v
[Jev]  supported / unsupported / declined ......... 1 request, ~200 ms, ~$0.000025
  |
  v
[code] supported and confidence >= 0.8 ?  -->  send
       otherwise                          -->  human review

整個流程的票務需要兩次 Jev 呼叫和一次 LLM 呼叫。總成本約為 $0.00015,總時間為 1.5 秒。兩次 Jev 呼叫使我們能使用廉價模型進行寫作,同時避免人工閱讀每一則模型回覆的需求。

根據自身資料設定閾值

無論閾值出現在程式碼的哪裡,請記住該數字並非無中生有。它來自你,代表你的政策選擇。Jev 的機率是 校準 的,基本上是大量預測後統計其正確率。因此若我們看到 0.8,表示大約 80% 的預測結果為正確。但並非每一次都是正確的,這只是一個平均值。請把本文中看到的 0.8 與 0.5 視為我們的數值,而非你自己的。

若想自行設定閾值,幾百個標記案例即可。將模型跑在所有案例上。對每個閾值,統計正確率。將準確率對照閾值繪製圖表。然後將閾值設在自動路徑與人類團隊回覆相符的點。修改條件後再檢查。別忘了。重新表述一個選項可能會以相當驚人的方式改變其他選項的機率。

先從一個 switch 語句開始

若你的程式碼庫有一個以 JSON.parse 結束的 LLM 呼叫,接著是 switch,這就是你的第一個 Jev 問題。將它替換進去,保留需要文字說明的分支使用 LLM,並測量兩者。

常見問題

什麼是 Jev,與 LLM 有何不同?

Jev 是 TypeSafe 的 System One 決策模型。它接受狀態加上已型別化的問題,問題可為固定集合中的選項、是非命題,或是你所描述層級的分數,並回傳機率而非文字。大型語言模型產生開放式文字,之後你需要解析並信任。

Jev 在分類上比小型 LLM 更便宜嗎?

於 2026 年 9 月 19 日在 OpenRouter 處理 60 張支援工單時,Jev 每千張工單成本 $0.0248,GPT Luna 為 $0.0921,Claude Opus 為 $2.88,意圖準確度分別為 59/60(Jev)、59/60(GPT Luna)以及 60/60(Claude Opus)。Jev 1.13 的輸入代幣每百萬計費 $0.042,輸出免費。價格可能隨時變動,請在預算前確認模型頁面。

如何將 Jev 與 LLM 一起使用?

兩種模式。首先,先路由:Jev 分類請求並回傳信心值。程式碼將該信心值與你設定的閾值比較,只有需要文字說明的請求才送往 LLM。其次,驗證後:先由 LLM 草擬,Jev 檢查草稿中每一項主張是否符合政策,僅在通過後才發送回覆。

Jev 使用哪個 OpenRouter 端點?

Jev 使用 Decisions API 端點,POST https://openrouter.ai/api/alpha/decisions 搭配模型 ID typesafe/jev-1.13,或在 TypeScript SDK 中使用 openrouter.alpha.decisions.create()。傳統 LLM 則使用 POST https://openrouter.ai/api/v1/chat/completions,或 openrouter.chat.send()。

我應該使用什麼信心閾值?

沒有一個適用於所有情況的通用數值。Jev 的機率在多個預測中經過校準,總體上可靠,但單一預測仍可能錯誤。因此,請從你自己的標記資料和錯誤成本中選擇閾值。例如,我們的路由規則是:意圖信心低於 0.8,或升級預測值 ≥ 0.5,則轉交人工。

來源:openrouter blog · openrouter.ai