orca-bench/orca-bench
Coding
Professional
ORCA-bench: root-cause analysis of oncall incidents — agents investigate telemetry from an instrumented microservices system to diagnose each incident; public split of 755 tasks across 77 incidents (245 easy, 264 medium, 246 hard).
Run this task
CLI:
inspect eval inspect_harbor/orca_bench --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import orca_bench
eval(orca_bench(), model="openai/gpt-5")Dataset information
| Harbor registry | orca-bench/orca-bench |
| Inspect task | orca_bench |
| Latest digest | sha256:3e53f8f8e64b58400549e793b280b59166a0dd64ccb656657c29c9f5f98b02c3 |
| Samples | 755 |
| Paper | arxiv |
| Source | https://github.com/orca-bench/ORCA-bench |
See Task Parameters for the parameter set shared across all Harbor tasks.