agentscope-ai/pawbench

Assistants
Coding

PawBench: A comprehensive agent benchmark with 150 tasks covering coding, data analysis, tool use, reasoning, safety, and multimodal capabilities.

← Back to Registry

Run this task

CLI:

inspect eval inspect_harbor/agentscope_ai_pawbench --model openai/gpt-5

Python:

from inspect_ai import eval
from inspect_harbor import agentscope_ai_pawbench

eval(agentscope_ai_pawbench(), model="openai/gpt-5")

Dataset information

Harbor registry agentscope-ai/pawbench
Inspect task agentscope_ai_pawbench
Latest digest sha256:a4197a2fe40bf045abad4ea2ab1037bf9c8bdfab01c378fdcc2b40242f5d63dd
Samples 150
Source https://github.com/agentscope-ai/PawBench

See Task Parameters for the parameter set shared across all Harbor tasks.