Harbor
Harbor is a AI agent skill for testing and qa, published by harbor-framework.
Harbor is a framework for evaluating arbitrary agents, building benchmarks, and conducting parallel experiments across thousands of environments.
Is it any good?
Effectiveness · benchmark pending
We haven't benchmarked Harbor yet, so its score is provisional — computed from safety, maintenance, and documentation only, with popularity given no weight. When we benchmark it, we measure whether it produces materially better output than a frontier model alone, and that result leads the score. How we measure.
About Harbor
Description · AI-summarised from the harbor-framework README
Harbor is a framework designed for evaluating and optimizing AI agents and language models. It enables users to evaluate arbitrary agents (Claude Code, OpenHands, Codex CLI, etc.), build and share custom benchmarks and environments, conduct large-scale parallel experiments through providers like Daytona and Modal, and generate rollouts for reinforcement learning optimization. Harbor is the official harness for Terminal-Bench-2.0 and supports multiple third-party benchmarks including SWE-Bench and Aider Polyglot.
Where it fits
Browse related skills
Install Harbor
Install
Install via agent tool
# install via agent tool
# see vendor docsInstall with your AI
Zero install · paste into ChatGPT, Claude, Cursor…
Install the "Harbor" skill into my project by following all instructions at https://skillsdirectory.co/s/harbor-framework-harbor
Your assistant reads /s/harbor-framework-harbor and follows it — fetching the files from the source, placing them, and confirming before it runs anything. No account needed. How this works.
Frequently asked
FAQ
What does Harbor do?
Harbor is a framework for evaluating arbitrary agents, building benchmarks, and conducting parallel experiments across thousands of environments.
How do I install Harbor?
Harbor is a AI agent skill. Primary install path: # install via agent tool # see vendor docs
What is the Skill Score for Harbor?
Harbor has a Skill Score of 6.8/10 (Solid). It is currently a provisional score — Harbor hasn't been benchmarked yet, so it weights safety 75/100 (40%), maintenance 51/100 (30%), and documentation 72/100 (30%). Popularity carries zero weight until effectiveness is measured. See methodology.
Is Harbor free?
Harbor is published under the Apache-2.0 license. See the vendor's source repository for any usage limits or paid tiers.
Related skills in Testing and QA
Related
LangSmith SDK
Client SDK for tracing, evaluating, and monitoring LLM apps in LangSmith — works with LangChain or standalone.
MCPJam Inspector
Debug, test, and evaluate MCP servers with full JSON-RPC tracing, multi-LLM chat, OAuth validation, and CI/CD integration.
aimock
One-package mock server for testing AI agents across 14+ LLM providers, vector databases, and AI protocols with zero dependencies.
Embed this score
For maintainers · always-current, links back
[](https://skillsdirectory.co/skills/harbor-framework-harbor)
<a href="https://skillsdirectory.co/skills/harbor-framework-harbor"><img src="https://skillsdirectory.co/badge/harbor-framework-harbor.svg" alt="Scored by skillsdirectory" height="28" /></a>