16 citations · 44 across the 35 of their papers we have counts for
12 papers · 1 filter
What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation
Kanghui Tian, Siyuan Liu, Tianxiang Jiang +8
More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a frozen copy of the base model scores the stu…
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current…
SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
Sheng Xia, Zhengqin Lai, Tianxiang Jiang +4
Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-tempo…
EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning
Zhitong Wang, Songze Li, Hao Peng +4
Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents. However, conventional RL methods for long-horizon agentic tasks…
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
Ziang Yan, Sheng Xia, Jiashuo Yu +10
Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant se…
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
Tianxiang Jiang, Linquan Wu, Sheng Xia +5
Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize intermediate future reasoning in…