Public benchmarks

Benchwright Registry

Open benchmark registry. Browse, run, and compare models — every public run lands here for everyone.

benchmarks
subsets
tasks
total runs
models
⌘K
back to registry
Built with the benchmark builder — 2 tasks · llm-judge grader. See how it was built →

Boot Zelda: Link's Awakening

Can the agent boot Link's Awakening (V1.2 U) from power-on and reach in-game play? Scored on wGameplayType phase reached.

2 samples 1 run · 1 impl impl-44dc5ddf6b9f
metricscore
What it measures
tool-usetextreward
Can the agent boot Link's Awakening (V1.2 U) from power-on and reach in-game play? Scored on wGameplayType phase reached.
How it's evaluated
Grader
impl-44dc5ddf6b9f
Task count
2
Category
tool-use
Avg duration
11m 51s
Avg tokens
164.9k in → 1.3k out
Platform cost
$0.0399
Platform cost is what Benchwright charges to run the eval (sandbox execution) — roughly model-independent. Model inference cost is paid to your provider and depends entirely on the model you pick (e.g. an Opus run costs far more than a small open model); see the leaderboard for per-model cost.
Top models
1
gpt-6-astra
11m 51s
0.830
Run it yourself
Pick the official fingerprint, give it a model, and benchwright will reproduce the exact evaluation. Public runs land back in the registry.
Download dataset (JSONL)2 tasks · 896 B