模型路由器基準測試
Model Router Benchmarks
summary_zh: OpenRouter 於 2026‑10‑02 推出 Model Router Benchmarks 頁面,讓使用者可比較 Auto Router、Free Models Router 等路由器在品質、速度與成本上的表現。
Today we’re introducing a new Model Router Benchmarks page for comparing performance across the growing number of router options. Every router was run through a set of benchmarks from varied domains to score them on quality, speed, and cost.
Individual models representing both best-in-class cost efficiency and intelligence are shown alongside routers as a baseline for comparison. Keep an eye on this page over time, as we’ll continually add new routers and update scores as the routers improve.

OpenRouter introduced two of the first model routers with our Auto Router and Free Models Router. They were created so you’d have a dependable option that stays up to date with the latest models, at your preferred level of cost. Since then we’ve continually improved them, including by using the wisdom of the market to inform routing decisions.
In recent months, many routers launched that use the differing strengths across the model landscape to optimize performance.
Cost differences: For example, comparing DeepSeek v4 Flash with GPT-6 Astra, GPT-6 Astra’s average price paid per token is over 48x higher, and it’s 21x more expensive to run an average 10 to 49 turn session in Codex. Routers can switch parts of a task to a lower-cost model to save on cost.
Task diversity: For example, the model that performs best on legal research may not be the same as the model that’s best for coding. Routers can switch between them to leverage each model’s strengths.
Session stage: Different parts of an agentic session may require different levels of intelligence. Routers can switch between models or reasoning efforts to adjust for parts of a session that require complex planning or mundane execution.
In practice, these techniques each have challenges that can sometimes lead to worse performance than using a single model directly:
- Switching between models means requests get more expensive while the input cache is rebuilt.
- Routers often don’t have enough information to assess task complexity by simply looking at the prompt.
- Heuristic signals for session stage may not agree with the model’s next step.
- The more processing you run on the routing decisions, the more latency is introduced.
We offer these Model Router Benchmarks to show which routers overcome these inherent challenges so you can decide when (and if) it’s the right time to put them to use. They’re designed to help you find the router and model combination that matches the performance characteristics you need.
Types of routers
The term “router” is used interchangeably to mean many different things. Let’s break it down.
Provider routing: Models are usually served by several inference providers, so a provider router picks the one that fits your preferences for price, speed, uptime, and data policy, then falls back to another if it fails. This is what OpenRouter has done since day one.
Model routing: Instead of choosing a single model directly, you can send a request to a router such as the Auto Router, and it decides which model should answer. These behave just like models, but may use one or more models to respond to your request.
The model routers we benchmarked take a few different approaches:
Routers that execute a blend of models: Unbiased’s Pareto and Sakana’s Fugu are inference providers that use a blend of models, sometimes with frontier escalation. They don’t disclose the models used or how often switching happens, and they bill at standardized per-token rates.
Routers that select a model: Auto Router and Jev Router select a model for each turn. The model selection is transparent, and you’re billed at the standard rate for whichever model was selected.
Routers that swap between a pre-defined set of models: NVIDIA’s Switchyard pairs a cheaper model with a stronger one and switches between them as an agent works through a task. Several pairs were benchmarked to show which models work best together, but you can run it with any pairing you want.
There are a couple of other styles, such as Pareto Code or the -latest model slugs, that alias to a single model. Pareto Code picks the most cost-optimal model for your desired level of intelligence, and the -latest slugs point to the most recent model within a family. Fusion is another style that runs the same request across a council of models and synthesizes the best answer. These styles weren’t benchmarked because they serve specialized purposes that warrant different comparison methods.
Router Index blends quality, speed, and cost
Benchmarks reveal performance across several dimensions, making it hard to tell which router performed best. The Router Index synthesizes quality, speed, and cost into a 0 to 10 score for each benchmark. By default, it weights quality (benchmark score) at 60%, time per task at 20%, and cost at 20%. Your priorities may differ, so you can drag a slider on the page to adjust the index weights.

Try any of these routers by simply swapping them in for your model
As with any benchmark, our Model Router Benchmarks are only representative of general tasks rather than your own work. All of the routers listed are available on OpenRouter. Swap them in for your model in any app, harness, or through our API.
來源:openrouter blog · openrouter.ai