Skip to content

Research1 min read

Harbor Adapters and Harbor-Index for Large-Scale Agent Evaluation

Introduces Harbor Adapters for evaluating agents across over 80 benchmarks and presents Harbor-Index, a curated set of challenging tasks for comprehensive assessment.

By OpenSmartRoute editorial · written through the router by llm-onprem

From arXiv cs.AI - “Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Harbor Adapters provide a unified infrastructure to evaluate arbitrary agents across more than 80 benchmarks, validated through code review and parity experiments.

A large-scale evaluation was conducted on 8 models across 54 benchmarks, using different harnesses, enabling broader analysis of agent capabilities and failure modes.

Harbor-Index is a curated set of 82 difficult, diverse, and high-quality tasks derived from the adapted benchmarks. It maintains evaluation challenge while remaining affordable, with no model exceeding a 30% pass rate.

All artifacts, including adapters, evaluation results, and Harbor-Index, are released as open-source to support reliable and comprehensive language-model agent evaluation.

Source: https://arxiv.org/abs/2609.04298

Published Sep 7, 2026 · updated Sep 7, 2026 · 99 words

Keep reading

Related posts

More in Research