3 papers
cs.CV2026
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
Minjoon Jung, Junbin Xiao, Junghyun Kim +2
Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introduce EgoExo-Con(sistency), a benc…
cs.RO2026
PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments
Kihyun Kim, Chaeyun Kim, Jongho Shin +4
Learning a good action embedding space is fundamental to scalable robot policy learning, yet existing methods treat action latents as task-specific intermediates rather than first-…
cs.RO2025
CLIP-RT: Learning Language-Conditioned Robotic Policies from Natural Language Supervision
Gi-Cheon Kang, Junghyun Kim, Kyuhwan Shim +2
Teaching robots desired skills in real-world environments remains challenging, especially for non-experts. A key bottleneck is that collecting robotic data often requires expertise…