2 papers
cs.AI2026
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin
Yuyao Sun, Tao Deng, Shuang Li +3
Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual…
cs.CL2026
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…