agentscope-ai/pawbench

Assistants
Coding

PawBench: A comprehensive agent benchmark with 150 tasks covering coding, data analysis, tool use, reasoning, safety, and multimodal capabilities.

← Back to Registry

Run this task

CLI:

inspect eval inspect_harbor/agentscope_ai_pawbench --model openai/gpt-5

Python:

from inspect_ai import eval
from inspect_harbor import agentscope_ai_pawbench

eval(agentscope_ai_pawbench(), model="openai/gpt-5")

Dataset information

Harbor registry agentscope-ai/pawbench
Inspect task agentscope_ai_pawbench
Latest digest sha256:b4e0ac2e4d5614fd46735f3ece1fb98207e02ea5818b07f9d5545f97fdde2bfe
Samples 150
Source https://github.com/agentscope-ai/PawBench

See Task Parameters for the parameter set shared across all Harbor tasks.