---
id: 20260918-T0-11
title: "Vercel Sandbox上线Harbor评测，单次试验跑独立微VM"
title_en: "Harbor Evals Now Run on Vercel Sandbox with Isolated MicroVMs"
url: https://ai.daily.yangsir.net/daily/20260918-T0-11
issue_date: 2026-09-18
publish_date: 2026-09-17T19:00:00.000Z
category: release
source_name: "Vercel Blog"
source_url: https://vercel.com/changelog/run-terminal-bench-and-other-harbor-evals-on-vercel-sandbox
---

# Vercel Sandbox上线Harbor评测，单次试验跑独立微VM

Vercel宣布可以在其Sandbox上运行Harbor评测框架。Harbor是Terminal-Bench背后的开源工具，其注册表还包含SWE-bench、tau3-bench和OSWorld等基准。使用时给harbor run传入--env vercel参数，每次试验都会在独立的Firecracker微VM中执行。这意味着开发者不需要自建隔离环境，就能跑多种Agent基准测试，适合需要频繁验证模型在真实终端任务中表现的团队。

## English Version

**Harbor Evals Now Run on Vercel Sandbox with Isolated MicroVMs**

Vercel now supports running Harbor evals on Vercel Sandbox. Harbor is the open-source harness behind Terminal-Bench, and its registry includes benchmarks like SWE-bench, tau3-bench, and OSWorld. Passing --env vercel to harbor run executes each trial in its own isolated Firecracker microVM. This lets developers run multiple agent benchmarks without building their own isolation infrastructure, useful for teams that frequently validate model performance on real terminal tasks.

---

**来源**：[Vercel Blog](https://vercel.com/changelog/run-terminal-bench-and-other-harbor-evals-on-vercel-sandbox)

**详情页**：https://ai.daily.yangsir.net/daily/20260918-T0-11

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*