blobfishai/hubbench
Assistants
Professional
HubBench: one oracle-proven, deterministically graded Blobfish benchmark family per Harbor Hub professional-domain cluster — 104 stateful multi-system employee-decision tasks over isolated SQLite worlds (HubScore, no LLM judge).
Run this task
CLI:
inspect eval inspect_harbor/blobfishai_hubbench --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import blobfishai_hubbench
eval(blobfishai_hubbench(), model="openai/gpt-5")Dataset information
| Harbor registry | blobfishai/hubbench |
| Inspect task | blobfishai_hubbench |
| Latest digest | sha256:9d61a3ef494915718a4ab16f2ad64f16e6193c281f3536288b8a95ccd59ece04 |
| Samples | 104 |
| Source | https://github.com/blobfishai/hub-agent-simulation |
See Task Parameters for the parameter set shared across all Harbor tasks.