5 papers
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
Siddhant Panpatil, Arth Singh, Mijin Koo +3
Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while…
PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments
Kihyun Kim, Chaeyun Kim, Jongho Shin +4
Learning a good action embedding space is fundamental to scalable robot policy learning, yet existing methods treat action latents as task-specific intermediates rather than first-…
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
Junhoo Lee, Mijin Koo, Nojun Kwak
Text-to-image models are commercially valuable assets often distributed under restrictive licenses, but such licenses are enforceable only when violations can be detected. Existing…
Targeted Data Protection for Diffusion Model by Matching Training Trajectory
Hojun Lee, Mijin Koo, Yeji Song +1
Recent advancements in diffusion models have made fine-tuning text-to-image models for personalization increasingly accessible, but have also raised significant concerns regarding…
Point-to-Point: Sparse Motion Guidance for Controllable Video Editing
Yeji Song, Jaehyun Lee, Mijin Koo +2
Accurately preserving motion while editing a subject remains a core challenge in video editing tasks. Existing methods often face a trade-off between edit and motion fidelity, as t…