11 papers
VGR: Visual Grounded Reasoning
Jiacong Wang, Zijian Kang, Haochen Wang +8
In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inherently suffers from language bias…
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
Fangxun Shu, Yongjie Ye, Yue Liao +6
We introduce SAIL-RL, a reinforcement learning (RL) post-training framework that enhances the reasoning capabilities of multimodal large language models (MLLMs) by teaching them wh…
GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation
Xiao Liang, Yunzhu Zhang, Linchao Zhu
Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches…
COLT: Enhancing Video Large Language Models with Continual Tool Usage
Yuyang Liu, Meng Cao, Xinyuan Shi +1
The success of Large Language Models (LLMs) has significantly propelled the research of video understanding. To harvest the benefits of well-trained expert models (i.e., tools), vi…
AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
Hongyuan Dong, Dingkang Yang, Xiao Liang +2
Learning rate is widely regarded as crucial for effective foundation model pretraining. Recent research explores and demonstrates the transferability of learning rate configuration…
Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks
Xiaoye Liang, Lai Jiang, Minglang Qiao +6
In recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, th…