atm-bench/atm-bench-hard-sgm

Assistants
Reasoning

ATM-Bench-Hard (SGM): 31 hard long-term personalized memory questions answered from schema-guided memory (SGM) files distilled from ~4 years of a user’s images, videos, and emails.

← Back to Registry

Run this task

CLI:

inspect eval inspect_harbor/atm_bench_hard_sgm --model openai/gpt-5

Python:

from inspect_ai import eval
from inspect_harbor import atm_bench_hard_sgm

eval(atm_bench_hard_sgm(), model="openai/gpt-5")

Dataset information

Harbor registry atm-bench/atm-bench-hard-sgm
Inspect task atm_bench_hard_sgm
Latest digest sha256:745358f9028fc12c5bb83ddb9d401e43791d0383e61a70b9f83d0d927af2d821
Samples 31
Paper arxiv
Source https://github.com/JingbiaoMei/ATM-Bench

See Task Parameters for the parameter set shared across all Harbor tasks.