6 papers · 1 filter
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
Siyuan Huang, Xiaoye Qu, Yafu Li +6
While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a "Visual Signal Dilution" phenomenon, where the accumul…
GEMS: Agent-Native Multimodal Generation with Memory and Skills
Zefeng He, Siyuan Huang, Xiaoye Qu +4
Recent multimodal generation models have achieved remarkable progress on general-purpose generation tasks, yet continue to struggle with complex instructions and specialized downst…
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
Zefeng He, Xiaoye Qu, Yafu Li +3
While recent Multimodal Large Language Models (MLLMs) have attained significant strides in multimodal reasoning, their reasoning processes remain predominantly text-centric, leadin…
VideoSSR: Video Self-Supervised Reinforcement Learning
Zefeng He, Xiaoye Qu, Yafu Li +3
Reinforcement Learning with Verifiable Rewards (RLVR) has substantially advanced the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, the rapi…
Spotlight on Token Perception for Multimodal Reinforcement Learning
Siyuan Huang, Xiaoye Qu, Yafu Li +4
While Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Vision-Language Models (LVLMs), most existing methods in multimodal rea…
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
Zefeng He, Xiaoye Qu, Yafu Li +3
While Large Vision-Language Models (LVLMs) have achieved substantial progress in video understanding, their application to long video reasoning is hindered by uniform frame samplin…