PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
PRISM evaluates programmatic video generation through code, focusing on whether models can produce spatially correct animated outputs rather than merely executable programs.
Highlights
- Provides 10,372 human-calibrated instruction-code pairs across English and Chinese prompts.
- Covers real-world knowledge visualization scenarios spanning hundreds of subject categories.
- Introduces a funnel-style evaluation framework for executability, spatial correctness, dynamic visual complexity, and temporal activity.
Citation
Qiran Zhang, Yuheng Wang, Runde Yang, Lin Wu, Jingru Fan, Shu Yao, Jie Zhang, Tianle Zhou, Huatao Li, Ruijie Shi, Yihan Li, and Chen Qian. (2026). "PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning." arXiv:2605.19382.