Harbor Adapters provide a unified infrastructure to evaluate arbitrary agents across more than 80 benchmarks, validated through code review and parity experiments.
A large-scale evaluation was conducted on 8 models across 54 benchmarks, using different harnesses, enabling broader analysis of agent capabilities and failure modes.
Harbor-Index is a curated set of 82 difficult, diverse, and high-quality tasks derived from the adapted benchmarks. It maintains evaluation challenge while remaining affordable, with no model exceeding a 30% pass rate.
All artifacts, including adapters, evaluation results, and Harbor-Index, are released as open-source to support reliable and comprehensive language-model agent evaluation.
Source: https://arxiv.org/abs/2609.04298