collaborators

5 papers

cs.CV2026

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

Shijie Li, Yilin Gao, Siyuan Yang +7

Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuanc…

cs.CV2026

From Priors to Perception: Grounding Video-LLMs in Physical Reality

Zicheng Zhao, Chaofan Gan, Shijie Li +1

While Video Large Language Models (Video-LLMs) excel in general understanding, they exhibit systematic deficits in fine-grained physical reasoning. Existing interventions not only…

cs.CV2026

VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding

Zhihao He, Tieyuan Chen, Kangyu Wang +6

Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this…

cs.RO2026

DV-VLN: Dual Verification for Reliable LLM-Based Vision-and-Language Navigation

Zijun Li, Shijie Li, Zhenxi Zhang +2

Vision-and-Language Navigation (VLN) requires an embodied agent to navigate in a complex 3D environment according to natural language instructions. Recent progress in large languag…

cs.CV2025

CogStream: Context-guided Streaming Video Question Answering

Zicheng Zhao, Kangyu Wang, Shijie Li +3

Despite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual…