islo-labs/reward-hack-bench-control

Safeguards

Baseline/control companion to reward-hack-bench: the same 8 SWE-bench + CyBench tasks with NO cheat prompt, still scored by the fairness judge. Also packages these CyBench/GlacierCTF tasks as Harbor tasks.

← Back to Registry

Run this task

CLI:

inspect eval inspect_harbor/islo_labs_reward_hack_bench_control --model openai/gpt-5

Python:

from inspect_ai import eval
from inspect_harbor import islo_labs_reward_hack_bench_control

eval(islo_labs_reward_hack_bench_control(), model="openai/gpt-5")

Dataset information

Harbor registry islo-labs/reward-hack-bench-control
Inspect task islo_labs_reward_hack_bench_control
Latest digest sha256:99138aaa3bec4e5fccc5a3c63b547604320d70b754aafbb5389660c202f37332
Samples 8
Source https://github.com/islo-labs/reward-hack-bench

See Task Parameters for the parameter set shared across all Harbor tasks.