15 papers
Learning Explicit Physical Parameter Control and Benchmarking for Video Generation
Yanxun Li, Hao Wen, Bingze Song +7
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Cu…
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
Ruofei Ju, Xinrui Wang, Xin Ding +12
Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since embodied environments vary across layouts,…
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
Mingzhe Huang, Weijun Wang, Xin Ding +7
In Vision-Language Models (VLMs), processing a massive number of visual tokens incurs prohibitive computational overhead. While recent training-aware pruning methods attempt to sel…
AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
Shizuo Tian, Hao Wen, Yuxuan Chen +6
The rapid development of mobile GUI agents has stimulated growing research interest in long-horizon task automation. However, building agents for these tasks faces a critical bottl…
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
Shidong Yang, Tongwen Huang, Hao Wen +3
Multimodal reward models are crucial for aligning multimodal large language models with human preferences. Recent works have incorporated reasoning capabilities into these models,…
From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
Yiqing Zhou, Yu Lei, Shuzheng Si +7
Managing extensive context remains a critical bottleneck for Large Language Models (LLMs), particularly in applications like long-document question answering and autonomous agents…