在 Vercel Sandbox 上執行 Terminal-Bench 及其他 Harbor 評估
Run Terminal-Bench and other Harbor evals on Vercel Sandbox
Vercel Sandbox 允許在隔離的 Firecracker 微VM 中執行 Harbor 評估,並可透過 AI Gateway 呼叫多個模型,提升測試效率與安全性。
You can now run Harbor evals on Vercel Sandbox.
Harbor is the open-source harness behind Terminal-Bench, whose registry includes many other benchmarks such as SWE-bench, tau3-bench and OSWorld. Pass --env vercel to harbor run and each trial executes in its own isolated Firecracker microVM, so you can parallelize far beyond what your local machine is capable of.
A task's network policy is enforced at the sandbox firewall, outside the VM. Optional credential injection attaches secrets to matching outbound requests at that firewall, so they never enter the sandbox.
Paired with AI Gateway, one AI_GATEWAY_API_KEY reaches hundreds of models from multiple providers, and benchmarking another model is the same command with a different --model:
uv tool install 'harbor[vercel]'
export VERCEL_TOKEN="<your-token>"
export AI_GATEWAY_API_KEY="<your-key>"
harbor run -d terminal-bench/terminal-bench-2-1 \
--agent fx \
--model vercel_ai_gateway/anthropic/claude-fable-5 \
--env vercel \
--n-concurrent 8
Swap --model to vercel_ai_gateway/openai/gpt-5.6-luna to run the same benchmark against an OpenAI model.
Requires Harbor 0.22.0 or later. Follow the step-by-step guide for setup, configuration, and troubleshooting. Learn more in the Sandbox documentation.
來源:Vercel Blog · vercel.com