3 papers
cs.CV2026
PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation
Wenbin Tan, Jiawen Lin, Fangyong Wang +4
3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentati…
cs.CV2026
Shot-Aware Frame Sampling for Video Understanding
Mengyu Zhao, Di Fu, Yongyu Xie +4
Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when…
cs.CV2025
Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
Jiahao Li, Yang Lu, Yachao Zhang +4
Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhan…