3 papers
cs.CV2025
FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation
Bingyu Li, Da Zhang, Zhiyuan Zhao +2
Open-vocabulary segmentation aims to identify and segment specific regions and objects based on text-based descriptions. A common solution is to leverage powerful vision-language m…
cs.CV2024
SignEye: Traffic Sign Interpretation from Vehicle First-Person View
Chuang Yang, Xu Han, Tao Han +5
Traffic signs play a key role in assisting autonomous driving systems (ADS) by enabling the assessment of vehicle behavior in compliance with traffic regulations and providing navi…
cs.CV2023
Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action Understanding
Shengkai Sun, Daizong Liu, Jianfeng Dong +5
Unsupervised pre-training has shown great success in skeleton-based action understanding recently. Existing works typically train separate modality-specific models, then integrate…