collaborators

9 papers

cs.CV2026

TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos

Yifan Mei, Qingling Shi, Changli Wu +3

Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perc…

cs.CV2026

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

Yuhui Zeng, Xinyu Mao, Xiaokun Liu +4

Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this inter…

cs.CV2026

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

Yuhui Zeng, Wang Chen, Jinfa Huang +5

Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. While recent efficient methods a…

cs.CV2026

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

Shuyan Ke, Yifan Mei, Changli Wu +4

Reasoning segmentation has recently expanded from ground-level scenes to remote-sensing imagery, yet UAV data poses distinct challenges, including oblique viewpoints, ultra-high re…

cs.CV2026

HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models

Yansong Guo, Chaoyang Zhu, Jiayi Ji +2

Video Large Language Models (VideoLLMs) have demonstrated impressive capabilities in video understanding, yet the massive number of input video tokens incurs a significant computat…

cs.CV2026

MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation

Changli Wu, Haodong Wang, Jiayi Ji +5

Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with o…