3 papers
cs.CV2026
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
Jiameng Li, Minye Wu, Jiezhang Cao +2
Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tokens while sparse sampling ri…
cs.CV2025
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
Zhaochong An, Guolei Sun, Yun Liu +5
Few-shot 3D point cloud segmentation (FS-PCS) aims at generalizing models to segment novel categories with minimal annotated support samples. While existing FS-PCS methods have sho…
cs.MM2024
Towards Open-Vocabulary Video Semantic Segmentation
Xinhao Li, Yun Liu, Guolei Sun +3
Semantic segmentation in videos has been a focal point of recent research. However, existing models encounter challenges when faced with unfamiliar categories. To address this, we…