17 papers
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
Sung-Hoon Yoon, Hoyong Kwon, Changgyoon Oh +1
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised model DINOv3 provides strong str…
Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training
Taewon Seo, Seonae Jeon, Giwon Lee +2
Accurate motion prediction of surrounding agents and safe motion planning are two closely coupled key tasks for social robot navigation in crowded environments. Deploying these sys…
Ego-Human Motion Prediction with 3D-Aware LLM
Yujin Bae, Jaewoo Jeong, Hyeonseong Kim +1
Anticipating human motion from an egocentric perspective is fundamental for proactive assistance in AR/VR, human-robot collaboration, and embodied AI. While recent works incorporat…
RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather
Heejun Park, Jaeseok Jeong, Kuk-Jin Yoon
Robust 3D object detection in adverse weather conditions is challenging due to sensor limitations. Although combining complementary modalities such as LiDAR and 4D RADAR has shown…
Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation
Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon
Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language…
Multimodal Distribution Matching for Vision-Language Dataset Distillation
Jongoh Jeong, Hoyong Kwon, Minseok Kim +1
Dataset distillation compresses large training sets into compact synthetic datasets while preserving downstream performance. As modern systems increasingly operate on paired vision…