A benchmark for evaluating multi-agent LLM coordination under OS-enforced Planner / Executor / Verifier role separation.
TeamBench measures the marginal contribution of each role in an LLM agent team. Roles run in separate Docker containers with disjoint filesystem mounts: the Planner reads the full spec but cannot edit, the Executor edits the workspace but cannot read the full spec, and the Verifier reads the spec and the read-only workspace but cannot modify either. Every task ships a deterministic shell-script grader.
931 evaluation instances · 19 categories · 5 ablation conditions · 27-configuration cross-provider grid · MIT-licensed.
With pip:
git clone https://github.com/ybkim95/TeamBench.git
cd TeamBench
pip install -e ".[all]" # add provider SDKs (anthropic, openai, google-genai)
docker compose build # OS-enforced role separation requires DockerOr with uv for a reproducible install from uv.lock:
git clone https://github.com/ybkim95/TeamBench.git
cd TeamBench
uv sync --all-extras
docker compose buildSet the providers you want to evaluate:
export ANTHROPIC_API_KEY=...
export OPENAI_API_KEY=...
export GEMINI_API_KEY=...A single task, one model, all 5 ablation conditions, seed 0:
python -m harness.ablation \
--model gemini-3-flash-preview \
--tasks DIST1_queue_race \
--seeds 0 \
--conditions oracle restricted full team_no_plan team_no_verify \
--output shared/runs/exampleThe output writes one score.json per (task, condition) under shared/runs/example/. Each file contains passed: true|false and a partial score in [0, 1] from the deterministic grader.
The full TeamBench-90 leaderboard sweep (all 5 conditions across the 90 stratified tasks):
python -m harness.ablation \
--model <your-model> \
--tasks $(jq -r '.tasks[].task_id' leaderboard/data/leaderboard_90_tasks.json) \
--seeds 0 \
--conditions oracle restricted full team_no_plan team_no_verify \
--output shared/ablation_results/lb90_<your-model>_seed0.jsonThe 90 task IDs are listed under the tasks[].task_id keys of leaderboard/data/leaderboard_90_tasks.json (a JSON object, not an array; jq -r '.tasks[].task_id' extracts them).
Aggregate scores and compute TNI:
python -m harness.compute_tni \
--ablation shared/ablation_results/lb90_<model>_seed0.json \
--output shared/ablation_results/tni_<model>.json
python -m harness.paper_tables \
--ablation shared/ablation_results/lb90_<model>_seed0.json \
--output-dir shared/paper/--model mock runs a deterministic stub that exercises the grading + sandboxing pipeline without calling any provider. Useful for CI and for sanity-checking new tasks:
python -m harness.ablation \
--model mock \
--tasks DIST1_queue_race \
--seeds 0 \
--conditions oracle \
--output /tmp/mock_runModels route by prefix; bring your own by implementing ToolCallAdapter:
from harness.agent_interface import ToolCallAdapter, AdapterResponse
class MyAdapter(ToolCallAdapter):
def generate_with_tools(self, messages, system_prompt, tools) -> AdapterResponse:
... # call your model, return text + tool_calls
def get_usage(self) -> dict:
return {"input_tokens": ..., "output_tokens": ..., "total_tokens": ...}Any OpenAI-compatible endpoint (vLLM, Ollama, Together AI, ...) works through the OpenAI adapter via --base-url.
tasks/<TASK_ID>/
├── spec.md # Full requirements (Planner-only)
├── brief.md # User-facing symptom (Executor)
├── workspace/ # Initial files
└── grade.sh # Deterministic grader, exits non-zero on failure
Optional generators/gen_<task_id>.py parameterizes the workspace from a seed (so seeds 5+ can be held out for the leaderboard refresh).
Smoke test the task:
python -m harness.ablation \
--model mock --tasks <TASK_ID> --seeds 0 1 2 \
--conditions oracle full --output /tmp/smoke_<TASK_ID>- Run the full sweep across the 5 conditions on the 90 leaderboard tasks (
leaderboard/data/leaderboard_90_tasks.json). - Open a PR adding
shared/ablation_results/lb90_<your-model>_seed0.json. - Maintainers manually re-run the deterministic graders to verify the submission before adding the model to the leaderboard; the leaderboard only ranks verified submissions.
Notes:
- One of the 90 leaderboard tasks,
GH120_redis-py_3863, is currently under re-curation. Runs against it return a 0 score until the spec is restored; the remaining 89 tasks are fully evaluable. Tracking issue will follow. - Graders run inside Docker sandboxes that write intermediate artifacts under
/tmpinside the container, so the host stays clean. When you use--model mockthe harness runs on the host instead, and a few grader-side files (such aspipe3_*.py,trap*_results.json) may appear under your host/tmp. They are safe to delete.
MIT (see LICENSE). Tasks adapted from public GitHub issue trackers and UCI datasets retain their respective upstream licenses.