most citedSeed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

6 citations · 6 across the 7 of their papers we have counts for

collaborators

8 papers

cs.CV2026

X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving

Bohao Zhao, Chengrui Wei, Guangfeng Jiang +17

Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive percept…

cs.CL2026

Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation

Ning Wu, Rui Liu, Xinkun Lin +5

Distilling demonstration effects into hidden-space interventions offers a lightweight alternative to full finetuning. However, existing multimodal variants are mostly evaluated on…

cs.CV2026

CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning

Zihan Lin, Songhe Deng, Shuwei He +6

Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, whereas segment-wise captions im…

cs.CV2026

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

Jianzhe Gao, Rui Liu, Yuxuan Xu +6

Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter percep…

cs.CV2026

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

Xudong Lu, Xueying Li, Annan Wang +8

We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline v…

cs.CV2026

Visual-Advantage On-Policy Distillation for Vision-Language Models

Ruiqi Liu, Xiaolei Lv, Gengsheng Li +8

On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-p…