collaborators

12 papers

cs.CV2026

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue

Yue Jiang, Xue Jiang, Lihua Zhang +6

Multimodal large language models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermined by hallucination snowball…

cs.RO2026

LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation

Xiangchen Wang, Weiye Zhu, Teng Wang +5

Recent navigation systems achieve strong benchmark results, yet real-world deployment often remains visibly stop-and-go. This bottleneck arises because the sense-inference-executio…

cs.CV2026

Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation

Jingnan Luo, Mingqi Gao, Jun Liu +2

The prosperity of Multimodal Large Language Models (MLLMs) has stimulated the demand for video reasoning segmentation, which aims to segment video objects based on human instructio…

cs.CV2025

R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios

Lu Zhu, Tiantian Geng, Yangye Chen +3

Recently, rapid advancements have been made in multimodal large language models (MLLMs), especially in video understanding tasks. However, current research focuses on simple video…

cs.CV2025

Video Understanding with Large Language Models: A Survey

Yolo Y. Tang, Jing Bi, Siting Xu +17

With the burgeoning growth of online video platforms and the escalating volume of video content, the demand for proficient video understanding tools has intensified markedly. Given…

cs.CV2025

UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization

Tiantian Geng, Teng Wang, Jinming Duan +4

Video event localization tasks include temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL). Existing methods tend to over-spe…