works on

From the 1 of 12 linked papers with an AI index.

activity
20242026
collaborators

12 papers

cs.CV2026

TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos

Yifan Mei, Qingling Shi, Changli Wu +3

Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perc…

cs.CV2026

GFR-SAM: Training-Free Referring Camouflaged Object Segmentation via Cross-Image Prompting

Yilong Yang, Jianxin Tian, Shengchuan Zhang +1

The paper introduces GFR‑SAM, a training‑free three‑stage framework that uses cross‑image prompting with SAM3 to segment camouflaged objects referenced by cues, employing exemplar‑…

cs.CV2026

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

Quanjian Song, Yefeng Shen, Mengting Chen +5

Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot support low-latency and interactiv…

cs.CV2026

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

Shaohui Dai, Yansong Qu, You Shen +2

Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual question answering, captioning,…

cs.CV2026

HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models

Yansong Guo, Chaoyang Zhu, Jiayi Ji +2

Video Large Language Models (VideoLLMs) have demonstrated impressive capabilities in video understanding, yet the massive number of input video tokens incurs a significant computat…

cs.CV2026

MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation

Changli Wu, Haodong Wang, Jiayi Ji +5

Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with o…