collaborators

19 papers

cs.CV2026

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

Yang Yang, Jiawei Chen, Tairan Chen +1

Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image…

cs.LG2026

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

Rosie Zhao, Anshul Shah, Xiaoyu Zhu +5

Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision-langua…

cs.CR2026

Red Teaming Large Reasoning Models

Jiawei Chen, Yang Yang, Chao Yu +6

Large Reasoning Models (LRMs) have emerged as a powerful advancement in multi-step reasoning tasks, offering enhanced transparency and logical consistency through explicit chains o…

cs.CV2026

AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

Zichuan Lin, Yicheng Liu, Yang Yang +2

Vision-Language Models (VLMs) have achieved remarkable success in visual question answering tasks, but their reliance on large numbers of visual tokens introduces significant compu…

cs.GR2026

RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer

Fangyu Du, Taiqing Li, Qian Qiao +7

Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve…

cs.CL2026

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…