5 papers
CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation
Sharath Girish, Tsai-Shien Chen, Zhikang Dong +4
Cinematic video depicts multiple subjects acting or interacting at specific moments, captured with deliberate camera movement, and stitched together by shot transitions. Together,…
Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
Tsai-Shien Chen, Aliaksandr Siarohin, Gordon Guocheng Qian +9
Visual concept personalization aims to transfer only specific image attributes, such as identity, expression, lighting, and style, into unseen contexts. However, existing methods r…
AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
Sharath Girish, Viacheslav Ivanov, Tsai-Shien Chen +3
Recent advances in subject-driven video generation with large diffusion models have enabled personalized content synthesis conditioned on user-provided subjects. However, existing…
Canvas-to-Image: Compositional Image Generation with Multimodal Controls
Yusuf Dalva, Guocheng Gordon Qian, Maya Goldenberg +5
While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when us…
LayerComposer: Multi-Human Personalized Generation via Layered Canvas
Guocheng Gordon Qian, Ruihang Zhang, Tsai-Shien Chen +10
Despite their impressive visual fidelity, existing personalized image generators lack interactive control over spatial composition and scale poorly to multiple humans. To address t…