AI product Open source
Harbor is a framework from the creators of Terminal-Bench for evaluating and optimizing AI agents and language models in container environments. It can run arbitrary agents, including Claude Code, OpenHands, and Codex CLI, against shared or custom benchmarks and environments such as Terminal-Bench, SWE-Bench, and Aider Polyglot.
Harbor launches benchmark runs locally with Docker or in parallel through cloud and sandbox providers, and can generate rollouts for reinforcement-learning optimization. It is distributed as a Python package installable with uv or pip and provides command-line tools for running datasets, selecting agents and models, and listing supported benchmarks.
1 use taken from transcripts — each links to the moment in the video.
Provides the evaluation framework referenced for coding-agent tasks.
1 in the library.