orca-bench/orca-bench

Coding
Professional

ORCA-bench: root-cause analysis of oncall incidents — agents investigate telemetry from an instrumented microservices system to diagnose each incident; public split of 755 tasks across 77 incidents (245 easy, 264 medium, 246 hard).

← Back to Registry

Run this task

CLI:

inspect eval inspect_harbor/orca_bench --model openai/gpt-5

Python:

from inspect_ai import eval
from inspect_harbor import orca_bench

eval(orca_bench(), model="openai/gpt-5")

Dataset information

Harbor registry orca-bench/orca-bench
Inspect task orca_bench
Latest digest sha256:3e53f8f8e64b58400549e793b280b59166a0dd64ccb656657c29c9f5f98b02c3
Samples 755
Paper arxiv
Source https://github.com/orca-bench/ORCA-bench

See Task Parameters for the parameter set shared across all Harbor tasks.