9 papers
ST-DiffEye: Diffusion-based Continuous Gaze Generation via Joint Scanpath-Trajectory Modeling
Brian Nlong Zhao, Ozgur Kara, Junho Kim +1
We study the problem of human gaze modeling, which aims to generate the gaze patterns a viewer produces while observing a visual stimulus. Gaze is primarily captured through two mo…
Narrative-Driven Paper-to-Slide Generation via ArcDeck
Tarik Can Ozden, Sachidanand VS, Furkan Horoz +3
We introduce ArcDeck, a multi-agent framework that formulates paper-to-slide generation as a structured narrative reconstruction task. Unlike existing methods that directly summari…
Grounding World Simulation Models in a Real-World Metropolis
Junyoung Seo, Hyunwook Choi, Minkyung Kwon +10
What if a world simulation model could render not an imagined environment but a city that actually exists? Prior generative world models synthesize visually plausible yet artificia…
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
Min-Seop Kwak, Junho Kim, Sangdoo Yun +4
We introduce a diffusion-based framework that performs aligned novel view image and geometry generation via a warping-and-inpainting methodology. Unlike prior methods that require…
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
Gayoung Lee, Junho Kim, Jin-Hwa Kim +1
Understanding reflection remains a long-standing challenge in 3D reconstruction due to the entanglement of appearance and geometry under view-dependent reflections. In this work, w…
StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
Jaeseok Jeong, Junho Kim, Gayoung Lee +2
In the domain of text-to-image generation, diffusion models have emerged as powerful tools. Recently, studies on visual prompting, where images are used as prompts, have enabled mo…