Enterprise-Bench/l1-l2-bench
Assistants
Professional
Enterprise-Bench (L1-L2): DevRev’s enterprise agent benchmark — answering L1/L2 support-tier questions (ticket SLAs, triage, customer-account data) over fragmented enterprise systems.
Run this task
CLI:
inspect eval inspect_harbor/enterprise_bench_l1_l2_bench --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import enterprise_bench_l1_l2_bench
eval(enterprise_bench_l1_l2_bench(), model="openai/gpt-5")Dataset information
| Harbor registry | Enterprise-Bench/l1-l2-bench |
| Inspect task | enterprise_bench_l1_l2_bench |
| Latest digest | sha256:834e40f54800287c913cebc9fa8e8273292940436168a56c1c5c72ff8a687e8d |
| Samples | 14 |
See Task Parameters for the parameter set shared across all Harbor tasks.