23 papers
VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
Ziyun Zeng, Zixuan Wang, Yongsheng Yu +2
Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide li…
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
Haoyu Chen, Kaichen Zhou, Hang Hua +11
The paper introduces MemoBench, a benchmark that tests video generation models' ability to remember and correctly update objects that disappear and later reappear in dynamically ch…
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
Niccolo Avogaro, Nayanika Debnath, Li Mi +6
Despite recent successes, test-time scaling -- i.e., dynamically expanding the token budget during inference as needed -- remains brittle for vision-language models (VLMs). Unstruc…
UXBench: Measuring the Actionability of LLM-Generated UX Critiques
Wenjie Wang, Yue Huang, Zipeng Ling +11
Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs. Yet no controlled benchmark measures…
Aligning Quantum Operators with Large Language Models
Rogerio Feris, Yunchao Liu, Pengyuan Li +2
Can Large Language Models (LLMs) understand and reason about quantum operators? Despite their remarkable capabilities in mathematics and symbolic reasoning, LLMs remain inherently…
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8
Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the genera…