4 papers
Segment-Aligned Policy Optimization for Multi-Modal Reasoning
Lei Gao, Zhuoming Li, Mengxi Jia +4
Existing reinforcement learning approaches for Large Language Models typically perform policy optimization at the granularity of individual tokens or entire response sequences. How…
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
Zhuoxu Huang, Mengxi Jia, Hao Sun +2
Reinforcement Learning with verifiable rewards (RLVR) has emerged as a primary learning paradigm for enhancing the reasoning capabilities of multi-modal large language models (MLLM…
Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding
Zhuoming Li, Aitong Liu, Mengxi Jia +5
Free-form gesture understanding is highly appealing for human-computer interaction, as it liberates users from the constraints of predefined gesture categories. However, the sole e…
Infinite Video Understanding
Dell Zhang, Xiangyu Chen, Jixiang Luo +6
The rapid advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have ushered in remarkable progress in video understanding. However, a fundamental ch…