9 papers
Text-based Tactile Graphics Generation for the Visually Impaired
Ruihan Gao, Joonghyuk Shin, Ava Pun +3
Tactile graphics are a primary medium for blind and low-vision (BLV) individuals to access non-textual information. However, they are difficult to scale or personalize. While recen…
ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation
Sanghyun Jo, Wooyeol Lee, Ziseok Lee +3
Recent open-weight text-to-image (T2I) diffusion models still struggle with multi-instance prompts, often omitting or merging instances and mixing semantics among similar objects.…
Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation
Youngjoong Kim, Duhoe Kim, Woosung Kim +1
Consistency models have been proposed for fast generative modeling, achieving results competitive with diffusion and flow models. However, these methods exhibit inherent instabilit…
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
Mingi Kwon, Joonghyuk Shin, Jaeseok Jung +2
The intrinsic link between facial motion and speech is often overlooked in generative modeling, where talking head synthesis and text-to-speech (TTS) are typically addressed as sep…
MotionStream: Real-Time Video Generation with Interactive Motion Controls
Joonghyuk Shin, Zhengqi Li, Richard Zhang +4
Current motion-conditioned video generation methods suffer from prohibitive latency (minutes per video) and non-causal processing that prevents real-time interaction. We present Mo…
Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild
Seunguk Do, Minwoo Huh, Joonghyuk Shin +1
Single-view 3D human reconstruction has achieved remarkable progress through the adoption of multi-view diffusion models, yet the recovered 3D humans often exhibit unnatural poses.…