collaborators

5 papers

cs.CL2026

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning

Yinan Zhou, Haokun Lin, Yichen Wu +7

Large multimodal models have achieved strong reasoning on complex visual tasks, but their inference efficiency is often restricted by long chains of thought. A promising solution i…

cs.HC2026

DataMagic: Transforming Tabular Data into Data Insight Video

Yupeng Xie, Chen Ma, Zhenyang Wang +6

Data videos integrate dynamic charts, voice narration, and synchronized animations to communicate data insights as temporal narratives, making them an effective medium for improvin…

cs.CV2026

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

Zijie Meng, Yufei Liu, Chengqian Ma +8

Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses resi…

cs.CV2025

DOGR: Towards Versatile Visual Document Grounding and Referring

Yinan Zhou, Yuxin Chen, Haokun Lin +6

With recent advances in Multimodal Large Language Models (MLLMs), grounding and referring capabilities have gained increasing attention for achieving detailed understanding and fle…

cs.IR2025

Scale Up Composed Image Retrieval Learning via Modification Text Generation

Yinan Zhou, Yaxiong Wang, Haokun Lin +3

Composed Image Retrieval (CIR) aims to search an image of interest using a combination of a reference image and modification text as the query. Despite recent advancements, this ta…