12 papers
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
Jooyeol Yun, Jintae Park, Hyesu Lim +3
Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering…
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Hoiyeong Jin, Hyojin Jang, Junha Hyung +6
Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D s…
Probability-Conserving Flow Guidance
Parsa Esmati, Junha Hyung, Amirhossein Dadashzadeh +2
Diffusion and flow-based generative models dominate visual synthesis, with guidance aligning samples to user input and improving perceptual quality. However, Classifier-Free Guidan…
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
Minho Park, Kinam Kim, Junha Hyung +5
Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet,…
Not the Example, but the Process: How Self-Generated Examples Enhance LLM Reasoning
Daehoon Gwak, Minseo Jung, Junwoo Park +4
Recent studies have shown that Large Language Models (LLMs) can improve their reasoning performance through self-generated few-shot examples, achieving results comparable to manual…
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
Min-Jung Kim, Jeongho Kim, Hoiyeong Jin +2
Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide creators with cinemati…