agentscope-ai/pawbench
Assistants
Coding
PawBench: A comprehensive agent benchmark with 150 tasks covering coding, data analysis, tool use, reasoning, safety, and multimodal capabilities.
Run this task
CLI:
inspect eval inspect_harbor/agentscope_ai_pawbench --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import agentscope_ai_pawbench
eval(agentscope_ai_pawbench(), model="openai/gpt-5")Dataset information
| Harbor registry | agentscope-ai/pawbench |
| Inspect task | agentscope_ai_pawbench |
| Latest digest | sha256:a4197a2fe40bf045abad4ea2ab1037bf9c8bdfab01c378fdcc2b40242f5d63dd |
| Samples | 150 |
| Source | https://github.com/agentscope-ai/PawBench |
See Task Parameters for the parameter set shared across all Harbor tasks.