Official toolkit and release artifacts for WRBench, a diagnostic benchmark for testing whether video world models preserve world state across camera motion.
WRBench provides three composable pieces:
- Compile: turn a camera instruction into a model-native payload and auditable sidecars. No GPU is required.
- Generate (optional): execute that payload through a configured model backend.
- Evaluate: score existing videos with the separable D1–D6 diagnostic contract.
The repository also bundles Natural-25 prompts and first frames, the frozen 23-model paper table, and the versioned paper reproducibility contract.
Python 3.10 or newer is required.
git clone https://github.com/JinPLu/WRBench.git
cd WRBench
pip install -e .Install optional prompt, first-frame, and development dependencies with pip install -e ".[all]". Model weights and GPU frameworks are intentionally not package dependencies.
List the available models and camera presets:
wrbench models
wrbench presetsCompile a go-and-return yaw request with a bundled Natural-25 first frame:
IMAGE="$(python - <<'PY'
from wrbench.datasets import natural25_first_frame_path
print(natural25_first_frame_path("bedroom_cat_bed_jump"))
PY
)"
wrbench generate \
--model wan22-fun-5b-cam \
--camera preset:yaw_LR \
--image "$IMAGE" \
--prompt "A house cat waits on the bedroom floor, facing the bed." \
--out out.mp4wrbench generate is compile-only by default: it writes the target trajectory, native payload, and metadata without loading a model. Use wrbench actions --camera "yaw:left:37@49" --fps 16 to inspect a custom camera script. See the camera-control reference and more compile examples.
First inspect the metric contract; this requires no model configuration:
wrbench eval contractFor D1–D6 scoring, copy wrbench.runtime.example.json to wrbench.runtime.json, fill in the scorer paths, and run:
wrbench eval --runtime-config wrbench.runtime.json run \
--manifest videos.json \
--out-dir eval_out \
--scorer-profile wrbench_default \
--sidecar-profile-gate mainScorer requirements and granular commands are documented in the evaluation guide. D1–D6 remain separate diagnostics; WRBench does not collapse them into one quality score.
Real generation is an explicit opt-in because each upstream model has its own environment and weights. Copy wrbench.runtime.example.json, configure a supported backend, then add --runtime-config /path/to/wrbench.runtime.json --no-dry-run to the compile command above. See the backend guide.
A generation-capable model needs one registry record, one adapter module, and one adapter import. Follow Adding a model, then validate it with:
wrbench doctor --model <model-key>| Dimension | What it measures | Applies when |
|---|---|---|
| D1-CamPrec | Requested-camera trajectory precision | An explicit target trajectory exists |
| D1-CamAlign | Prompt-camera intent alignment | Camera control is prompt/API based |
| D2 | Visual integrity | All videos |
| D3 | Visible spatial consistency | All videos |
| D4 | Visible state consistency | All videos |
| D5 | Re-observation spatial consistency | The relevant content returns into view |
| D6 | Re-observation event-state consistency | The relevant content returns into view |
Rows without re-observation support are excluded from D5/D6 rather than counted as failures. The exact contract lives in src/wrbench/eval/contract/.
| Artifact | Source |
|---|---|
| Paper | arXiv:2606.20545 |
| All releases | Hugging Face collection |
| Natural-25 | dataset · bundled data |
| Frozen paper prompts, first frames, camera scopes, and source mappings | paper_main_20260608 |
| Results | rolling dataset · bundled frozen 23-model tables |
| Videos and per-video scores | dataset |
| Human annotations | dataset |
| Leaderboard | Space |
The 9,600-row paper table and its provenance are frozen. Corrections or newly verified assets belong to a separately versioned rolling release; they do not silently rewrite the paper aggregates. Exact prompt, first-frame, TV2V-source, and HyDRA policies are owned by the linked release contract rather than repeated here.
- Camera-control grammar
- Evaluation and scorer configuration
- Generation backends
- Adding a model
- Prompt generation
- First-frame generation
@article{wrbench2026,
title = {Current World Models Lack a Persistent State Core},
author = {Jinpeng Lu and Dexu Zhu and Haoyuan Shi and Linghan Cai and Guo Tang and Yinda Chen and Jie Cao and Duyu Tang and Yi Zhang and Yong Dai and Xiaozhu Ju},
journal = {arXiv preprint arXiv:2606.20545},
year = {2026},
url = {https://arxiv.org/abs/2606.20545},
}Apache 2.0 — see LICENSE.
