harbor-index/harbor-index-1.0
Assistants
Reasoning
Harbor Index is an 82-task benchmark for agentic evaluation. It was distilled from more than 6,000 candidate tasks through repeated model runs, automated broken-task identification, human audits, and reward hacking supervision.
Run this task
CLI:
inspect eval inspect_harbor/harbor_index_1_0 --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import harbor_index_1_0
eval(harbor_index_1_0(), model="openai/gpt-5")Dataset information
| Harbor registry | harbor-index/harbor-index-1.0 |
| Inspect task | harbor_index_1_0 |
| Latest digest | sha256:9d4514cb93f6fafd9cf8ff352c784495ab675176c7f09671db523bd19b663584 |
| Samples | 82 |
See Task Parameters for the parameter set shared across all Harbor tasks.