atm-bench/atm-bench-hard-sgm
Assistants
Reasoning
ATM-Bench-Hard (SGM): 31 hard long-term personalized memory questions answered from schema-guided memory (SGM) files distilled from ~4 years of a user’s images, videos, and emails.
Run this task
CLI:
inspect eval inspect_harbor/atm_bench_hard_sgm --model openai/gpt-5Python:
from inspect_ai import eval
from inspect_harbor import atm_bench_hard_sgm
eval(atm_bench_hard_sgm(), model="openai/gpt-5")Dataset information
| Harbor registry | atm-bench/atm-bench-hard-sgm |
| Inspect task | atm_bench_hard_sgm |
| Latest digest | sha256:745358f9028fc12c5bb83ddb9d401e43791d0383e61a70b9f83d0d927af2d821 |
| Samples | 31 |
| Paper | arxiv |
| Source | https://github.com/JingbiaoMei/ATM-Bench |
See Task Parameters for the parameter set shared across all Harbor tasks.