5 papers
AdaCodec: A Predictive Visual Code for Video MLLMs
Haowen Hou, Zhen Huang, Zheming Liang +8
Video is temporally redundant: adjacent frames usually share most objects, background, and layout. Yet existing video multimodal large language models (video MLLMs) usually encode…
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
Chenglin Li, Qianglong Chen, Feng Han +6
Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning over uniformly sampled frames,…
EtCon: Edit-then-Consolidate for Reliable Knowledge Editing
Ruilin Li, Yibin Wang, Wenhong Zhu +5
Knowledge editing aims to update specific facts in large language models (LLMs) without full retraining. Prior efforts sought to tune the knowledge layers of LLMs, achieving improv…
VideoPro: Adaptive Program Reasoning for Long Video Understanding
Chenglin Li, Feng Han, Yikun Wang +9
Large language models (LLMs) have shown promise in generating program workflows for visual tasks. However, previous approaches often rely on closed-source models, lack systematic r…
RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
Jinghao Zhang, Naishan Zheng, Ruilin Li +4
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a promising framework for improving reasoning abilities in Large Language Models (LLMs). However, poli…