19 papers
Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs
Yang Yang, Jiawei Chen, Tairan Chen +1
Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image…
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
Rosie Zhao, Anshul Shah, Xiaoyu Zhu +5
Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision-langua…
Red Teaming Large Reasoning Models
Jiawei Chen, Yang Yang, Chao Yu +6
Large Reasoning Models (LRMs) have emerged as a powerful advancement in multi-step reasoning tasks, offering enhanced transparency and logical consistency through explicit chains o…
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
Zichuan Lin, Yicheng Liu, Yang Yang +2
Vision-Language Models (VLMs) have achieved remarkable success in visual question answering tasks, but their reliance on large numbers of visual tokens introduces significant compu…
RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
Fangyu Du, Taiqing Li, Qian Qiao +7
Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…