blobfishai/hubbench

Assistants
Professional

HubBench: one oracle-proven, deterministically graded Blobfish benchmark family per Harbor Hub professional-domain cluster — 104 stateful multi-system employee-decision tasks over isolated SQLite worlds (HubScore, no LLM judge).

← Back to Registry

Run this task

CLI:

inspect eval inspect_harbor/blobfishai_hubbench --model openai/gpt-5

Python:

from inspect_ai import eval
from inspect_harbor import blobfishai_hubbench

eval(blobfishai_hubbench(), model="openai/gpt-5")

Dataset information

Harbor registry blobfishai/hubbench
Inspect task blobfishai_hubbench
Latest digest sha256:9d61a3ef494915718a4ab16f2ad64f16e6193c281f3536288b8a95ccd59ece04
Samples 104
Source https://github.com/blobfishai/hub-agent-simulation

See Task Parameters for the parameter set shared across all Harbor tasks.