blobfishai/devopsbench-100
Coding
Professional
100 deterministic long-horizon DevOps/SRE agent tasks over an executable NovaCart world with 32-67 call reference trajectories and state-diff vcode verifiers.
Run this task
CLI:
inspect eval inspect_harbor/blobfishai_devopsbench_100 --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import blobfishai_devopsbench_100
eval(blobfishai_devopsbench_100(), model="openai/gpt-5")Dataset information
| Harbor registry | blobfishai/devopsbench-100 |
| Inspect task | blobfishai_devopsbench_100 |
| Latest digest | sha256:1efd654cd3248d225964702c8b4c6f7c9f3909a557b7a9c64eee91642e12efce |
| Samples | 100 |
| Source | https://github.com/blobfishai/software-devops |
See Task Parameters for the parameter set shared across all Harbor tasks.