8 papers
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
Rui Zhao, Kaiming Yang, Jifeng Zhu +6
Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domain. A natural question follow…
Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation
Sike Xiang, Shuang Chen, Kevin Qinghong Lin +4
Routine clinical check-up reports combine laboratory measurements, physiological assessments, imaging findings and visually structured information, but rarely tell patients what to…
Reasoning Compression with Mixed-Policy Distillation
Han Yang, Mingyan Wu, Bailan He +4
Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive token usage and high inferenc…
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
Jiaqi Wang, Weijia Wu, Yi Zhan +6
With AI-generated videos increasingly indistinguishable from reality, current benchmarks primarily focus on broad semantic alignment and basic physical consistency, offering limite…
Code2World: A GUI World Model via Renderable Code Generation
Yuhao Zheng, Li'an Zhong, Yi Wang +6
Autonomous GUI agents interact with environments by perceiving interfaces and executing actions. As a virtual sandbox, the GUI World model empowers agents with human-like foresight…