agentscope-ai/pawbench
Assistants
Coding
PawBench: A comprehensive agent benchmark with 150 tasks covering coding, data analysis, tool use, reasoning, safety, and multimodal capabilities.
Run this task
CLI:
inspect eval inspect_harbor/agentscope_ai_pawbench --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import agentscope_ai_pawbench
eval(agentscope_ai_pawbench(), model="openai/gpt-5")Dataset information
| Harbor registry | agentscope-ai/pawbench |
| Inspect task | agentscope_ai_pawbench |
| Latest digest | sha256:b4e0ac2e4d5614fd46735f3ece1fb98207e02ea5818b07f9d5545f97fdde2bfe |
| Samples | 150 |
| Source | https://github.com/agentscope-ai/PawBench |
See Task Parameters for the parameter set shared across all Harbor tasks.