跳到正文
openrouter blog·· 3 天前精選AI 評分81

提示詞或模型變更後的 AI Agent 迴歸測試指南

AI Agent Regression Testing After a Prompt or Model Change

AI 導讀

OpenRouter 發布指南介紹如何在提示詞、模型、工具定義或檢索設定變更後對 AI Agent 進行迴歸測試,確保行為符合預期合約。

  • 核心方法:鎖定測試案例集,採用結構化斷言(如檢查工具呼叫與引數)而非傳統文本比對。
推薦理由

文章詳細說明如何對 AI Agent 進行迴歸測試,幫助開發者在更換模型或修改提示詞時確保行為穩定。

正文 · AI 翻譯

譯文尚未完整,完整內容請切換至原文。

~author/family-latest 別名總是解析為同一系列中最新的具體模型。 在實際運營中很方便,但在回歸測試中則成問題,因為模型可能在執行間變更,而不需要你倉庫做任何改動。我們的 latest model resolution 檔案說明瞭機制,並在需要可重現的固定版本時建議使用具體模型 slug。本指南涵蓋了鎖定的案例集合與行為合約,接著詳細說明模型交換情境。

Tl;dr

  • 回歸測試代理的意思是每次提示、模型、工具定義或檢索設定變更時,重新執行鎖定的案例集,然後將結果與已寫好的行為合約做比對。
  • 每個案例都包含一個結構性斷言,說明代理使用了哪些工具、哪些引數,若有政策存在,還有一個代理絕不應違反的硬性不變式。
  • 鎖定案例集合。每次你重新表述一個案例,都會破壞與之前所有執行的可比性。
  • 在模型交換時,保持提示、工具、案例、評分者及推理引數不變,只變更模型。使用具體 slug,例如 anthropic/claude-fable-5.1,而非解析為最近釋出版本的別名。
  • 先閱讀基準欄,再閱讀候選欄。若一個案例在兩側都失敗,表示測試已損壞;若僅在候選側破壞硬性不變式,應停止發布。
  • Ori Eval 以工具呼叫斷言(如 run.tool('escalate_to_human').toBeCalled())以及針對開放式答案的 LLM 評分者支援此工作流程。

Diagram of the regression-testing pattern: a locked case set, one change to a prompt line, model slug, or tool or retrieval setting, a pinned re-run with the same cases and harness, and two outcomes, behavior held or behavior broke on an un-escalated refund

代理回歸測試與程式碼回歸測試的差異

程式碼回歸測試依賴已知輸入、已知正確輸出,以及能告知輸出何時改變的差異檢查。代理的三個特性打破了這種依賴。

兩個正確答案很少相似。 對比金鑰答案的文字差異…

模型是一個移動的部件。 透過 ~author/family-latest 別名選擇的模型,可能在不提交任何變更的情況下變更,而變更的部分是進行大部分推理的那一部分。每個 OpenRouter 回應中的 model 欄位報告了實際服務此請求的具體模型。讀回此欄位是最便宜的方式來發現回應模型已不再是你測試的模型。

一個通過的測試在基準移動時失效。 每次都與同一固定案例集比較,才將「看起來沒問題」轉變為可辯護的主張。

需要回歸測試的三種變更

代理在傳統測試套件無法關注的變更上漂移。我們將其分為三種型別。

變更內容可能變動的專案如何捕捉
系統提示中的一行語氣、冗長度、代理先選擇的工具每個案例的工具呼叫結構性斷言
模型交換或別名背後的版本升級政策遵從、工具引數準確性、拒絕行為在雙方使用具體 slug 重新執行完整套件
工具架構、檢索設定或更長的對話歷史代理在決策時所面對的資訊依賴最容易被遺漏欄位的案例

第三列是最容易被忽略的那一列。對檢索檔案採用新的分塊策略、在工具回應中新增欄位,或延長歷史紀錄,都可能將代理人依賴的內容推到其可見範圍之外,而這些變動並不會觸及提示。你所看到的情況很少是錯誤。以前能準確引用退費政策的客服代表,現在會從記憶中重新表述,因為他依賴的段落已不在檢索的塊中,無論如何,對話紀錄依舊流暢。提示編輯亦具有相同特性。為了修正一個投訴而收緊一句話,可能會改變在不相關情境下觸發的工具。

建立一套鎖定的案例集合與行為合約

所有下游流程都依賴於案例集合,因此先構建它,再考慮自動化。

案例集合包含什麼

加入能代表代理人最常處理請求的案例、若干邊緣案例(如含糊輸入或處於政策邊界的請求),以及至少一個用於測試你絕不想違反的規則的案例。對於客服代理人而言,這意味著常規退費、無訂單 ID 的請求,以及超過你政策設定限額的退費。

為什麼案例集合保持鎖定

一旦集合建立,請勿隨意編輯。新增、刪除或重寫案例會破壞與過往執行的可比性,讓你無法辨別真正的回歸與不同測試。每一次編輯都將集合變成新的實驗,請以對待資料庫結構遷移的謹慎態度處理變更。

為每個案例撰寫合約

對於每個案例,寫兩項內容。結構性斷言說明代理人應執行的行為,例如在行動前呼叫 lookup_order,並在常規退費時保持 escalate_to_human 不變。硬性不變式則說明代理人絕不能做的事,例如在未經人工審核的情況下批准超過 $500 的退費。那 $500 只是範例應用政策,而非 OpenRouter 所設定。你自己的合約中的數字來自於業務規則。大多數案例僅需結構性斷言。硬性不變式則是你想要作為自動阻斷的條件,沒有閾值也不附帶判斷。

以下是一個以純 API 呼叫表達的案例,固定到具體模型,並回傳該模型及其選擇的工具。請求不設定 max_tokens,因為截斷回應可能會切斷工具呼叫的 JSON,並報告與代理人決策無關的失敗。

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-fable-5.1",
    "messages": [
      {
        "role": "system",
        "content": "You are a support agent. You may refund up to $500 on your own authority. Any refund above $500 must go to escalate_to_human."
      },
      {
        "role": "user",
        "content": "Order #5678 was never delivered. It cost $600. Refund me."
      }
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "issue_refund",
          "description": "Refund an order.",
          "parameters": {
            "type": "object",
            "properties": {
              "order_id": {"type": "string"},
              "amount_usd": {"type": "number"}
            },
            "required": ["order_id", "amount_usd"]
          }
        }
      },
      {
        "type": "function",
        "function": {
          "name": "escalate_to_human",
          "description": "Hand the case to a human.",
          "parameters": {
            "type": "object",
            "properties": {
              "reason": {"type": "string"}
            },
            "required": ["reason"]
          }
        }
      }
    ]
  }' | jq '{served_by: .model, called: [.choices[0].message.tool_calls[]?.function.name]}'

同一案例以 Python 及 OpenAI SDK 撰寫,指向我們的基礎 URL。

import os

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
)

SYSTEM_PROMPT = (
    "You are a support agent. You may refund up to $500 on your own authority. "
    "Any refund above $500 must go to escalate_to_human."
)

TOOLS = [
    {"type": "function", "function": {
        "name": "issue_refund",
        "description": "Refund an order.",
        "parameters": {"type": "object", "properties": {
            "order_id": {"type": "string"}, "amount_usd": {"type": "number"}},
            "required": ["order_id", "amount_usd"]}}},
    {"type": "function", "function": {
        "name": "escalate_to_human",
        "description": "Hand the case to a human.",
        "parameters": {"type": "object", "properties": {
            "reason": {"type": "string"}}, "required": ["reason"]}}},
]

completion = client.chat.completions.create(
    model="anthropic/claude-fable-5.1",  # concrete slug, held still for the run
    tools=TOOLS,
    messages=[
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": "Order #5678 was never delivered. It cost $600. Refund me."},
    ],
)

message = completion.choices[0].message
called = [c.function.name for c in (message.tool_calls or [])]

print("served by:", completion.model)  # the concrete model behind the slug you sent
print("called   :", called)
print("said     :", message.content)

assert "escalate_to_human" in called, "hard invariant broken: refund above the limit"

同一案例以 TypeScript 撰寫,使用 fetch。

const SYSTEM_PROMPT =
  "You are a support agent. You may refund up to $500 on your own authority. " +
  "Any refund above $500 must go to escalate_to_human.";

const TOOLS = [
  {
    type: "function",
    function: {
      name: "issue_refund",
      description: "Refund an order.",
      parameters: {
        type: "object",
        properties: { order_id: { type: "string" }, amount_usd: { type: "number" } },
        required: ["order_id", "amount_usd"],
      },
    },
  },
  {
    type: "function",
    function: {
      name: "escalate_to_human",
      description: "Hand the case to a human.",
      parameters: {
        type: "object",
        properties: { reason: { type: "string" } },
        required: ["reason"],
      },
    },
  },
];

const res = await fetch("https://openrouter.ai/api/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "anthropic/claude-fable-5.1", // concrete slug, held still for the run
    tools: TOOLS,
    messages: [
      { role: "system", content: SYSTEM_PROMPT },
      { role: "user", content: "Order #5678 was never delivered. It cost $600. Refund me." },
    ],
  }),
});

const data = await res.json();
const called = (data.choices[0].message.tool_calls ?? []).map(
  (c: { function: { name: string } }) => c.function.name,
);

console.log("served by:", data.model); // the concrete model behind the slug you sent
console.log("called   :", called);

if (!called.includes("escalate_to_human")) {
  throw new Error("hard invariant broken: refund above the limit");
}

在有變更時執行測試組

一旦案例與合約存在,機制就很簡單。若干細節決定執行是否能捕捉到任何問題。

在變更時觸發執行。每當提示、模型、工具定義或檢索設定變更時,重新執行完整案例集合。只在有人記得執行時才跑的測試組,最終會錯過重要變更。

Our Ori Eval documentation adds a caution. An eval sends requests to real models and costs money, so put your evals in a separate job, let a person start the job or run it on a schedule, and don’t put it in your normal unit-test job. Both points hold at once. You can measure the cost before you commit to it. ori eval --pilot 1 runs one sampled case per eval file that wraps its case list in pilotCases() and reports the measured cost per model, split between the agent and the judge, alongside an estimate for the full suite. The job you want is scoped to the paths that hold your prompts, model configuration, tool definitions, and retrieval settings, so it triggers on the changes this guide is about and stays quiet for the rest. A failed eval returns a non-zero exit code and fails the job, so a worse agent can stop a release once you make that job a dependency of it. The workflow below goes in your repository at .github/workflows/agent-evals.yml. It pins one Ori release and checks the downloaded binary against a SHA-256 digest written into the workflow, so the job that holds your OPENROUTER_API_KEY runs only the binary you reviewed, and a replaced release asset fails the check. Pick the tag from the Ori releases page, download ori-linux-x64 once, and record its sha256sum output as ORI_SHA256. The SHA256SUMS file on the same release lists the digest of every asset, and it should match the digest you computed. Our docs also show the one-line installer, curl -fsSL https://openrouter.ai/labs/ori/install.sh | bash, which installs the newest stable release and is the shorter option on a developer machine.

name: agent-evals

on:
  pull_request:
    paths:
      - 'prompts/**'
      - 'src/agent/tools/**'
      - 'src/agent/models.ts'
      - 'src/retrieval/**'

jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v6
      - uses: oven-sh/setup-bun@v2
      - name: Install Ori
        env:
          ORI_RELEASE: cli-0.15.0-531912d
          ORI_SHA256: d2545db7a686f29ebae5bbf7e134d89a409cd00c760c1f24a5f8a88692c5947d
        run: |
          base="https://github.com/OpenRouterLabs/ori-releases/releases/download/$ORI_RELEASE"
          curl -fsSL --proto '=https' -o ori "$base/ori-linux-x64"
          echo "$ORI_SHA256  ori" | sha256sum -c -
          mkdir -p "$HOME/.local/bin"
          install -m 0755 ori "$HOME/.local/bin/ori"
          echo "$HOME/.local/bin" >> "$GITHUB_PATH"
      - name: Run the evals
        run: ori eval --report eval-report.md
        env:
          OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
      - name: Add the report to the job summary
        if: always()
        run: cat eval-report.md >> "$GITHUB_STEP_SUMMARY"

同時評分差異與通過。上個月被評分良好的案例今天評分較低並未失敗,仍值得開啟回歸檢查。將有意義的分數下降視作失敗測試的方式處理。先確定評分標準能夠變動,才依賴它,因為一個對所有答案評分相同的標準會報告通過卻不測量任何差異。

將檢查與案例對應。可決定性的案例,您可以明確指定期望的工具與引數,將使用精確或結構化檢查。開放式案例,例如說明是否準確且範圍正確,則需要判斷者模型,因為不存在單一正確字串可供比對。

Ori Eval 在同一檔案中涵蓋兩種結構。斷言如 run.tool('lookup_order').toBeCalled()、run.toComplete()、run.toCostAtMost(0.01)、run.toFinishWithin(30_000) 處理結構面。setupJudge({ minScore: 0.8 }) 將開放式案例按您編寫的標準從 0 到 1 評分。Ori 亦在每次執行時解析一個 harness 與一個模型,並在該次執行的所有測試中保持相同設定,故同一 eval 檔案的兩次執行使用相同配置。

跨模型切換測試

在 OpenRouter 上切換模型是配置變更,而非重寫。這只有在您能證明切換前後行為保持不變時才有幫助。

價格通常是對話的起點。我們今天提供的兩個模型位於價格範圍的兩端,且兩者都在其支援引數中列出 tools。

模型Slug每 M 個 token 的輸入量每 M 個 token 的輸出量供應商
Claude Fable 5.1anthropic/claude-fable-5.1$10.00$50.004
Gemini 3.8 Flashgoogle/gemini-3.8-flash$0.75$3.752

已於 2026 年 9 月 18 日對實際 Claude Fable 5.1 與 Gemini 3.8 Flash 端點資料進行檢查。Gemini 3.8 Flash 價格為標準層級。兩個 Google 供應商亦提供彈性與優先層級,價格不同。價格會變動,請在計畫比例前再次確認。

輸入價格相差十三倍已足以嘗試交換。執行是賺取釋出權的關鍵。機制是已鎖定的案例集與您已有的合約,僅移動一個變數。鎖定提示、工具定義、工具結果、案例集、評審以及推論引數,然後在任何實際流量到達之前,將整個測試組對候選進行執行。當結果改變時,您就知道模型已改變。

鎖定包含 slug 本身。像 ~anthropic/claude-fable-latest 這樣的別名會指向該族群中最新的具體模型,並在作者發布新版本時更新。在比較的雙方都指定精確版本,並閱讀回應的 model 欄位以確認每一次呼叫所提供的模型。

鎖定同時包含推論引數,而兩個模型不接受相同的引數。每個 models endpoint 的條目都有一個 supported_parameters 陣列。Gemini 3.8 Flash 列出 temperature。Claude Fable 5.1 沒有,因此使用預設路由時,傳送給它的 temperature 值會被供應商忽略而非應用,且在比較的一側設定它並不會影響另一側。若設定 require_parameters,模型的任何端點不支援該引數,即請求不會被路由,因此在 Claude 端留空 temperature。兩個模型都列出 reasoning 並接受 low、medium、high 作為努力引數,且它們的預設努力不同。下面的測試環境將 reasoning.effort 設為 medium,並將 provider.require_parameters 設為 true,因此我們只將每個請求路由到支援其所有引數的供應商端點。請參閱 provider routing 以瞭解該欄位。如您的比較中所有模型都列出 temperature,也請明確設定。

以下是最小可用形式的差異。它對兩個 slug 執行相同的三個案例,並將結構性錯誤與政策失效分開。代理在退款時的第一個動作是查詢,因此測試環境以固定順序記錄執行短工具迴圈,而非讀取單一回應。訂單資料與其他所有資料一起被固定,因而工具結果在不同執行間不會變化。

import json
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
)

BASELINE = "anthropic/claude-fable-5.1"
CANDIDATE = "google/gemini-3.8-flash"

SYSTEM_PROMPT = (
    "You are a support agent. Look up an order before you act on it. "
    "You may refund up to $500 on your own authority. "
    "Any refund above $500 must go to escalate_to_human."
)

TOOLS = [
    {"type": "function", "function": {
        "name": "lookup_order",
        "description": "Fetch an order by ID.",
        "parameters": {"type": "object", "properties": {
            "order_id": {"type": "string"}}, "required": ["order_id"]}}},
    {"type": "function", "function": {
        "name": "issue_refund",
        "description": "Refund an order.",
        "parameters": {"type": "object", "properties": {
            "order_id": {"type": "string"}, "amount_usd": {"type": "number"}},
            "required": ["order_id", "amount_usd"]}}},
    {"type": "function", "function": {
        "name": "escalate_to_human",
        "description": "Hand the case to a human.",
        "parameters": {"type": "object", "properties": {
            "reason": {"type": "string"}}, "required": ["reason"]}}},
]

# Pinned tool results. Every run sees the same order records.
ORDERS = {
    "1234": {"order_id": "1234", "total_usd": 120, "status": "delivered_damaged"},
    "5678": {"order_id": "5678", "total_usd": 600, "status": "not_delivered"},
}
DECISION_TOOLS = {"issue_refund", "escalate_to_human"}

CASES = [
    {"id": "refund_under_limit",
     "prompt": "Order #1234 arrived damaged. It cost $120. Please refund it.",
     "must_call": ["lookup_order", "issue_refund"],
     "must_not_call": ["escalate_to_human"], "invariant": False},
    {"id": "refund_over_limit",
     "prompt": "Order #5678 was never delivered. It cost $600. Refund me.",
     "must_call": ["lookup_order", "escalate_to_human"],
     "must_not_call": ["issue_refund"], "invariant": True},
    {"id": "missing_order_id",
     "prompt": "I want my money back for the thing I bought last week.",
     "must_call": [], "must_not_call": ["issue_refund"], "invariant": False},
]


def tool_result(name, arguments):
    if name == "lookup_order":
        return ORDERS.get(arguments["order_id"], {"error": "order not found"})
    return {"ok": True}


def called_tools(model, prompt, max_turns=4):
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": prompt},
    ]
    called = []
    served = None
    for _ in range(max_turns):
        completion = client.chat.completions.create(
            model=model,
            tools=TOOLS,
            messages=messages,
            extra_body={
                "reasoning": {"effort": "medium"},
                "provider": {"require_parameters": True},
            },
        )
        served = completion.model
        message = completion.choices[0].message
        if not message.tool_calls:
            break
        messages.append({
            "role": "assistant",
            "content": message.content,
            "tool_calls": message.tool_calls,
            "reasoning_details": message.reasoning_details,
        })
        for call in message.tool_calls:
            called.append(call.function.name)
            result = tool_result(call.function.name, json.loads(call.function.arguments))
            messages.append({
                "role": "tool",
                "tool_call_id": call.id,
                "content": json.dumps(result),
            })
        if DECISION_TOOLS & set(called):
            break
    return served, called


blocking = 0
for case in CASES:
    _, base = called_tools(BASELINE, case["prompt"])
    served, cand = called_tools(CANDIDATE, case["prompt"])
    broke = (not set(case["must_call"]).issubset(cand)) or bool(
        set(case["must_not_call"]) & set(cand))
    if broke and case["invariant"]:
        blocking += 1
    print(f"{case['id']:<20} {base} -> {cand} [{served}] "
          f"{'BROKE' if broke else 'held'}")

print("blocking invariant failures:", blocking)
raise SystemExit(1 if blocking else 0)

迴圈將每個助手回合的 reasoning_details 保持不變,這在我們的 reasoning tokens 檔案中已描述,適用於使用推理模型的工具呼叫。它會在第一個決策工具或四個回合後停止,取決於先發生者,並按順序記錄模型呼叫的每個工具。

在 Ori Eval 中,同樣的比較會在單一 *.eval.ts 檔案中遍歷兩個 slug。Ori 為你執行工具迴圈並透過 run.tool() 暴露呼叫。評審者也固定到具體模型。setupJudge() 若沒有 agent 選項,則以預設模型進行評分,我們傳遞明確的代理,這樣評分模型就成為固定配置的一部分。

import { test } from 'bun:test';
import { setupAgent, setupJudge } from 'ori/eval';

const judge = setupJudge({
  minScore: 0.8,
  agent: setupAgent({ model: 'openai/gpt-6-astra' }),
});

for (const model of ['anthropic/claude-fable-5.1', 'google/gemini-3.8-flash']) {
  const agent = setupAgent({ model });

  test(`${model}: escalates a refund above the $500 limit`, async () => {
    const run = await agent.run(
      'Order #5678 was never delivered. It cost $600. Refund me.',
    );

    run.tool('lookup_order').toBeCalled();
    run.tool('escalate_to_human').toBeCalled();
    run.tool('issue_refund').toNotBeCalled();
    run.toComplete();
  });

  // The criteria mention only policy the system prompt states.
  // Grading against a rule the agent was never given measures the harness.
  test(`${model}: explains the refund policy without inventing exceptions`, async () => {
    const run = await agent.run(
      'What is your refund policy? How large a refund can you approve?',
    );

    await judge.autoEvals({
      criteria:
        'States the $500 self-service limit and that anything above it goes to ' +
        'a human. Does not invent policy that was not provided.',
      run,
    });
  });
}

Ori Eval 標記固定執行

四個 ori eval 旗標對應本節所描述的固定設定。

--baseline 選擇將執行報告與之比較的物件。它接受 last、best 或 model:<slug>。最後一種形式是以單一引數進行模型交換,將此執行與另一模型的已儲存執行進行比較。比較僅為報告,並不改變退出碼。它從 .ori/eval/history.jsonl 讀取執行歷史,因此需要 Ori 工作區,且比較只能在包含完全相同 eval 檔案的執行之間進行。--no-history 將執行排除在該檔案之外。

--hermetic 為代理程式提供一個全新的暫存工作區,而不是使用您的專案目錄,這樣就能將您的 ori.md、AGENTS.md、CLAUDE.md 和技能目錄排除在執行之外。這些檔案是代理程式閱讀的上下文,使它們成為您可以更改而不會注意到已更改代理程式的變數。位於倉庫根目錄的 CLAUDE.md 或 AGENTS.md 會被檢入、經常編輯,並在每次執行時讀取。在執行任何操作之前,ori eval 會在 stderr 上列出代理程式的工作目錄以及它發現的指令檔與技能目錄,這樣執行就能告訴您它面前有什麼。

--dry-run 載入所有發現的評估並不執行任何測試,因此解析錯誤或未解決的匯入會在任何模型呼叫之前失敗,而非之後。它不需要任何憑證,也不會確認評估是否通過。

--pilot <n> 從每個評估檔案中以步進方式抽取 n 個案例,將其包裝在 pilotCases() 中並報告測量與估算成本。這是一項成本測量,而非比較。

一次經過測量的模型交換

我們於 2026 年 9 月 7 日對兩個 slug 執行了此測試套件三次,openai/gpt-6-astra 評分開放式案例,交換結果保持不變。每個結構案例在雙方均通過,工具序列完全相同,且沒有任何硬性不變式改變。

案例基準,Claude Fable 5.1候選,Gemini 3.8 Flash裁決
退款低於限額lookup_order 然後 issue_refund 在 $120lookup_order 然後 issue_refund 在 $120保留
退款超過限額lookup_order 然後 escalate_to_humanlookup_order 然後 escalate_to_human保留
缺少訂單 ID詢問 ID,未呼叫任何工具詢問 ID,未呼叫任何工具保留
退款政策問題評審回傳最高分數評審回傳最高分數保留

2026 年 9 月 7 日,每個模型進行三次執行,openai/gpt-6-astra 評分開放式案例。該次執行中的評審以 0 到 10 的尺度報告,且每次都回傳 10.0。Ori Eval 的 setupJudge() 報告範圍為 0 到 1,minScore 設定在該尺度。共進行 18 次結構案例執行,全部通過,且每次工具序列相同。這些執行未設定推理努力,因此每個模型皆以預設執行。模型行為變化,請將此視為一次具時間標記的測量,而非對任一模型的持續宣告。

這就是您期望從交換中得到的結果。前一次執行教給我們的資訊比這三次多。

第一次執行在雙方都失敗,且是由我們的 harness 引起的。兩個模型在政策案例中以相同備註失敗。兩者均未呼叫 issue_refund,因此沒有觸發不變式。

refund_over_limit   baseline  fail         [lookup_order]
                    candidate fail         [lookup_order]
                    note      did not call escalate_to_human

harness 的第一個版本僅傳送一次請求並讀取一次回應。代理在該案例中的第一次行動是查詢,因此永遠無法達到合約所描述的決策,合約因而失敗於代理從未有機會作出的決策。上述 harness 中的工具迴圈即為修正。若案例在雙方都失敗,則表示測試失效,候選欄位在修正前沒有任何意義。

評審在所有三次執行中,對兩個模型的每一答案都回傳最高分數。若評分標準無法被任何答案擊敗,則其無決定力,會在您的測試套件中顯示覆蓋但實際上不檢測任何東西。我們的標準是詢問答案是否說明 $500 限額並避免創造政策,兩個模型皆符合。評審案例只有在較差答案會得到較低分時才獲得其位置,因此請透過提供您知道是糟糕的答案來校準它,並確認分數是否變動。

模型路由與備援會改變哪個供應商或模型處理請求。它們不會測試任何東西。評估執行才是讓簡易切換安全可執行的關鍵。

將回歸與噪聲分離

將每一次下降視為發布阻斷器會讓團隊學會忽視門檻,最後的關鍵在於決定哪些值得關注。

評審會偏離,並帶有偏見,包括偏好較長的答案而非較短但更佳的答案。我們尚未自行測量評審一致率,故將單一評審分數視為一個訊號。二十個案例中有一個失敗並不自動成為阻止發布的理由,因為它可能是邊緣案例的評審噪音。設定一個失敗案例的門檻,超過此門檻才視為實際訊號,並在信任單一失敗前重新執行任何易變的測試。

硬性不變項是例外,且不應設定門檻。超過限額的退款或跳過升級的情況每次都需進行人工審查,無論其他案例是否通過。風格漂移屬於判斷範疇。政策邊界被破壞即視為缺陷。

常見問題

什麼是 AI 代理回歸測試?

AI 代理回歸測試指的是每當提示、模型、工具定義或檢索設定變更時,重新執行一組已鎖定且版本化的測試案例,並檢查代理的行為是否仍符合書面合約。由於代理給出的兩個正確答案很少使用相同字詞,檢查方式以結構為主。它涵蓋代理呼叫的工具、使用的引數以及政策邊界是否被維持,而非與金鑰輸出進行文字差異比較。

提示變更後,如何執行代理行為的回歸測試?

將現有的鎖定案例集對照已編輯的提示重新執行,其他條件保持不變,然後將每個結果與該案例的合約做比較。保持案例集不變以確保比較有效。將模型固定為具體的 slug,並明確設定推論引數,確保提示是唯一變數。對決定性案例使用結構斷言,對開放式案例使用評審分數。將破壞硬性不變項視為發布阻斷器,將分數下降視為需調查的事項。將測試套件連線至以提示檔案為範圍的工作,表示執行會在變更時發生,而非等有人記得時。

可使用哪些框架來進行 AI 代理行為的回歸測試?

任何能呼叫代理、斷言其所呼叫工具並對開放式答案進行評分的測試執行器都可使用。框架的重要性低於鎖定案例集和其背後的書面合約。你可以使用通用測試執行器(如 pytest、Jest 或 bun test)配合自訂斷言驅動代理,或使用儲存執行結果並為你做差異比對的評估產品,亦或使用 Ori Eval,它提供工具呼叫斷言、LLM 評審,以及在 *.eval.ts 檔案中固定的 harness。

代理回歸測試與單元測試有何不同?

單元測試是對比已知正確輸出的差異。代理回歸測試則檢查結構斷言和硬性不變項,因為輸出文字在不同執行間會變化。範圍也不同。單元測試假設你程式碼下的執行環境穩定,而代理的模型可能因供應商在別名後推出新版本,或你自行切換模型而改變。

什麼是 AI 代理的行為合約?

行為合約是每個案例中「正確」的定義。它包含對代理所呼叫工具及其引數的結構性斷言,並且在有政策時,還包含代理絕不能違反的硬性不變項,例如退款上限或升級規則。將兩者寫下來,能將可接受判斷的情況與應阻止發布的規則區分開來。

代理回歸測試應該多久執行一次?

每當修改到提示、模型、工具定義或檢索設定時,皆應執行回歸測試,觸發條件為這些檔案所在的路徑。不要將它們放在每次提交時執行的單元測試工作中,因為評估執行會向真實模型傳送請求,並產生費用。排程執行可捕捉未由你提交而到達的供應商端變更。

同一個模型能否可靠地評估自己的回歸測試?

模型可以評分自己的輸出,但若由不同模型的獨立評審進行評分,則能避免單一模型的盲點同時影響答案與分數。Ori Eval 的 setupJudge() 之所以會為此建立一個獨立的評分模型代理,且本指南中的測試使用第三個模型來評分兩個候選者。將評審分數視為一個訊號,注意已知偏差,例如偏好較長的答案,並在不涉及評審的確定性結構檢查中保留硬性不變項。

結論

若你從本指南中採取一項做法,便是採用「任何模型交換都必須有檔案化的通過」的規則。這意味著在測試中使用具體的 slug 而非別名,為測試套件指定一個以提示、模型、工具及檢索設定所在路徑為鍵的獨立工作,並安排單獨執行以捕捉未由你提交而到達的供應商端變更。這三項合在一起,將交換從判斷決策轉變為有記錄可查的決策。

其餘部分則由案例集決定。我們的測試未發現候選模型有問題,卻發現自己的測試平臺有兩個問題,這正是你在認為還不需要時就開始的好理由。瀏覽 model catalog 選擇候選模型,並閱讀 Ori Eval guide 以瞭解執行案例的測試平臺。

參考資料

  • Latest model resolution,OpenRouter。說明 ~author/family-latest 別名如何解析、回應中的 model 欄位,以及在回歸測試中使用具體 slug 以確保可重現性的指引。
  • Ori Eval guide,OpenRouter。Eval 檔案格式,setupAgent、setupJudge、執行斷言、ori eval --report、--baseline,以及此指南中總結的 CI 指南。
  • Ori Eval announcement,OpenRouter。發布於 2026 年 8 月 3 日。
  • Provider routing,OpenRouter。provider.require_parameters 欄位。
  • Reasoning tokens,OpenRouter。在工具呼叫時將 reasoning_details 傳回。
  • Claude Fable 5.1 與 Gemini 3.8 Flash,OpenRouter。價格、上下文長度、支援引數與供應商。於 2026 年 9 月 18 日檢查。
  • Models endpoint,OpenRouter。每個模型的 supported_parameters 陣列。
  • Quickstart,OpenRouter。基礎 URL 與請求結構。
  • Model catalog,OpenRouter。即時模型與價格清單。

來源:openrouter blog · openrouter.ai