7 papers
Aligned but Stereotypical? How System Prompts Shape Demographic Bias in LLM-Based Text-to-Image Models
NaHyeon Park, Na Min An, Kunhee Kim +3
Text-to-image (T2I) systems increasingly rely on Large Language Model (LLM)-based text conditioning to interpret and expand user prompts. While this improves prompt understanding a…
Grounding Driving VLA via Inverse Kinematics
Junsung Park, Hyunjung Shim
Existing Driving VLAs predict trajectories while largely ignoring their visual tokens -- a phenomenon we trace not to insufficient training but to a structurally ill-posed task for…
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
NaHyeon Park, Kunhee Kim, Hyunjung Shim
In this paper, we introduce TextBoost, an efficient one-shot personalization approach for text-to-image diffusion models. Traditional personalization methods typically involve fine…
SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals
Soyeon Yoon, Chang Wook Seo, Hyunjung Shim
Learning dense correspondences across deformable 3D shapes remains a long-standing challenge due to structural variability, non-isometric deformation, and inconsistent topology. Ex…
Representation Alignment for Just Image Transformers is not Easier than You Think
Jaeyo Shin, Jiwook Kim, Hyunjung Shim
Representation Alignment (REPA) has emerged as a simple way to accelerate Diffusion Transformers training in latent space. At the same time, pixel-space diffusion transformers such…
Directional Textual Inversion for Personalized Text-to-Image Generation
Kunhee Kim, NaHyeon Park, Kibeom Hong +1
Textual Inversion (TI) is an efficient approach to text-to-image personalization but often fails on complex prompts. We trace these failures to embedding norm inflation: learned to…