13 papers
Toward Efficient Agents: Memory, Tool learning, and Planning
Xiaofang Yang, Lijun Li, Heng Zhou +12
Recent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents has continued to improve, efficiency, whi…
TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment
Jiaming Li, Chenyu Zhu, Nanxi Yi +9
Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from significant reward hacking, which…
Rethinking Video-Language Model from the Language Input Perspective
Xiang Fang, Wanlong Fang, Changshuo Wang +2
Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although…
Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs
Xiang Fang, Wanlong Fang, Changshuo Wang +4
Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and…
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
Siyuan Huang, Xiaoye Qu, Yafu Li +6
While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a "Visual Signal Dilution" phenomenon, where the accumul…
Spotlight on Token Perception for Multimodal Reinforcement Learning
Siyuan Huang, Xiaoye Qu, Yafu Li +4
While Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Vision-Language Models (LVLMs), most existing methods in multimodal rea…