2 papers
cs.CV2025
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
Juan Yeo, Soonwoo Cha, Jiwoo Song +2
Vision-language models such as CLIP have recently propelled open-vocabulary dense prediction tasks by enabling recognition of a broad range of visual concepts. However, CLIP still…
cs.AI2025
Object-Centric World Model for Language-Guided Manipulation
Youngjoon Jeong, Junha Chun, Soonwoo Cha +1
A world model is essential for an agent to predict the future and plan in domains such as autonomous driving and robotics. To achieve this, recent advancements have focused on vide…