3 papers
cs.CV2026
SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering
Wenli Li, Kai Zhao, Haoran Jiang +3
Vision-language models (VLMs) have been widely adopted for 3D question answering (3D QA). In typical pipelines, visual tokens extracted from multiple viewpoints are concatenated wi…
cs.CV2025
Motion-Aware Vision-Reference Alignment for Referring Multi-Object Tracking
Weiyi Lv, Ning Zhang, Hanyang Sun +5
Referring Multi-Object Tracking (RMOT) extends conventional multi-object tracking (MOT) by introducing natural language references for multi-modal fusion tracking. RMOT benchmarks…
cs.CV2024
Task Consistent Prototype Learning for Incremental Few-shot Semantic Segmentation
Wenbo Xu, Yanan Wu, Haoran Jiang +3
Incremental Few-Shot Semantic Segmentation (iFSS) tackles a task that requires a model to continually expand its segmentation capability on novel classes using only a few annotated…