activity
20242026
most citedERNIE 5.0 Technical Report

2 citations · 2 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL20262 cited

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…

cs.CV2026

Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model

Jiedong Zhuang, Lu Lu, Ming Dai +4

Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual to…

cs.CV2025

Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models

Ying Peng, Hongsen Ye, Changxin Huang +3

Vision Transformers (ViTs) have achieved strong performance in video action recognition, but their high computational cost limits their practicality. Lightweight CNNs are more effi…

cs.HC2025

GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting

Kaichun Yang, Jian Chen

We present a quantitative evaluation to understand the effect of zero-shot large-language model (LLMs) and prompting uses on chart reading tasks. We asked LLMs to answer 107 visual…

cs.CL2025

KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues

Junhao Chen, Yu Huang, Siyuan Li +7

Multi-Turn Long-Form Question Answering (MT-LFQA) is a key application paradigm of Large Language Models (LLMs) in knowledge-intensive domains. However, existing benchmarks are lim…

cs.CL2025

Rectified Sparse Attention

Yutao Sun, Tianzhu Ye, Li Dong +6

Efficient long-sequence generation is a critical challenge for Large Language Models. While recent sparse decoding methods improve efficiency, they suffer from KV cache misalignmen…