Documentation
WebAgentBench Documentation
WebAgentBench is a research benchmark for evaluating web agents in seeded, stateful web environments. These docs cover the benchmark contract, scoring model, runtime architecture, and cognitive primitive taxonomy used to analyze agent behavior.
Scoring
How evaluation works — base score, penalties, trajectory modifier, and the full formula.
Cognitive Primitives
The 12-primitive taxonomy for diagnosing where and why web agents fail.
Architecture
Runtime layers, seeded sessions, audit-backed scoring, and trajectory tooling.
Benchmark
Task structure, seeded environments, version history, and how to run it.