5 papers
UI-Venus-2 Technical Report
Venus Team, Zhuohan Cai, Haoxing Chen +28
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remai…
VISTA: View-Consistent Self-Verified Training for GUI Grounding
Xinyu Qiu, Yunzhu Zhang, Heng Jia +3
When applying Group Relative Policy Optimization (GRPO) for GUI Grounding, rollouts are sampled from a single screenshot view; groups often become either all failures on difficult…
GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation
Xiao Liang, Yunzhu Zhang, Linchao Zhu
Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches…
MVP: Multiple View Prediction Improves GUI Grounding
Yunzhu Zhang, Zeyu Pan, Zhengwen Zeng +3
GUI grounding, which translates natural language instructions into precise pixel coordinates, is essential for developing practical GUI agents. However, we observe that existing gr…
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
Yunzhu Zhang, Yu Lu, Tianyi Wang +3
Long-form video understanding poses a significant challenge for video large language models (VideoLLMs) due to prohibitively high computational and memory demands. In this paper, w…