5 papers
PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning
Youngjoon Jeong, Jihwan Yu, Minsoo Jo +2
Latent action pretraining learns representations of visual change from pairs of observations, but existing methods typically encode each transition as a single unstructured represe…
Uncovering Vulnerability of Vision-Language-Action Models under Joint-Level Physical Faults
Minsoo Jo, Taeju Kwon, Junha Chun +2
Deploying Vision-Language-Action (VLA) models in real robotic systems requires robustness not only to semantic and perceptual variations, but also to embodiment-side faults that ch…
Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection
Yewon Han, Yumin Seol, EunGyung Kong +2
Existing jailbreak defence frameworks for Large Vision-Language Models often suffer from a safety utility tradeoff, where strengthening safety inadvertently degrades performance on…
Angular Gradient Sign Method: Uncovering Vulnerabilities in Hyperbolic Networks
Minsoo Jo, Dongyoon Yang, Taesup Kim
Adversarial examples in neural networks have been extensively studied in Euclidean geometry, but recent advances in \textit{hyperbolic networks} call for a reevaluation of attack s…
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
Keonwoo Kim, Yeongjae Cho, Taebaek Hwang +2
Recent research has demonstrated that Large Language Models (LLMs) are not limited to text-only tasks but can also function as multimodal models across various modalities, includin…