red-hat-ai/SWE-benchify-hard
Coding
Hard subset of SWE-bench-style Go issue-resolution tasks generated by Red Hat’s SWE-benchify pipeline from cloud-native repos (openshift, operator-framework, grpc-go, opentelemetry-go).
Run this task
CLI:
inspect eval inspect_harbor/red_hat_ai_swe_benchify_hard --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import red_hat_ai_swe_benchify_hard
eval(red_hat_ai_swe_benchify_hard(), model="openai/gpt-5")Dataset information
| Harbor registry | red-hat-ai/SWE-benchify-hard |
| Inspect task | red_hat_ai_swe_benchify_hard |
| Latest digest | sha256:6cffe94f4c17dd421fda7bb0b284d3ed1cc5250512d8f7b8e7120770bc220375 |
| Samples | 284 |
| Source | https://github.com/Red-Hat-AI-Innovation-Team/SWE-benchify |
See Task Parameters for the parameter set shared across all Harbor tasks.