activity
20242026
collaborators

13 papers

cs.CV2026

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra +4

Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as…

cs.CV2026

Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing

Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen +3

Video-to-video diffusion models achieve impressive single-turn editing performance, but practical editing workflows are inherently iterative. When edits are applied sequentially, e…

cs.CV2025

V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties

Ye Fang, Tong Wu, Valentin Deschaintre +6

Large-scale video generation models have shown remarkable potential in modeling photorealistic appearance and lighting interactions in real-world scenes. However, a closed-loop fra…

cs.CV2025

ReasonX: MLLM-Guided Intrinsic Image Decomposition

Alara Dirik, Tuanfeng Wang, Duygu Ceylan +2

Intrinsic image decomposition aims to separate images into physical components such as albedo, depth, normals, and illumination. While recent diffusion- and transformer-based model…

cs.CV2025

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders

Dohun Lee, Hyeonho Jeong, Jiwook Kim +2

Video diffusion models have advanced rapidly in the recent years as a result of series of architectural innovations (e.g., diffusion transformers) and use of novel training objecti…

cs.GR2025

PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling

Alara Dirik, Tuanfeng Wang, Duygu Ceylan +2

We present PRISM, a unified framework that enables multiple image generation and editing tasks in a single foundational model. Starting from a pre-trained text-to-image diffusion m…