collaborators

6 papers

cs.CV2026

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

Minjun Kang, Inkyu Shin, Taeyeop Lee +3

Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video d…

cs.MM2026

Inference-Time Scaling for Joint Audio-Video Generation

Jaemin Jung, Kyeongha Rho, Inkyu Shin +1

Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint au…

cs.CV2025

Drag4D: Align Your Motion with Text-Driven 3D Scene Generation

Minjun Kang, Inkyu Shin, Taeyeop Lee +2

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories f…

eess.AS2025

SCORE: Scaling audio generation using Standardized COmposite REwards

Jaemin Jung, Jaehun Kim, Inkyu Shin +1

The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advanc…

cs.CV2025

Deeply Supervised Flow-Based Generative Models

Inkyu Shin, Chenglin Yang, Liang-Chieh Chen

Flow based generative models have charted an impressive path across multiple visual generation tasks by adhering to a simple principle: learning velocity representations of a linea…

cs.CV2025

Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting

Inkyu Shin, Qihang Yu, Xiaohui Shen +3

Recent advancements in zero-shot video diffusion models have shown promise for text-driven video editing, but challenges remain in achieving high temporal consistency. To address t…