5 papers
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
Taewon Kang, Divya Kothandaraman, Ming C. Lin
Generating coherent long-form video sequences from discrete text prompts remains challenging due to difficulties in maintaining temporal coherence, semantic consistency, and scene-…
Low-Bitrate Video Compression through Semantic-Conditioned Diffusion
Lingdong Wang, Guan-Ming Su, Divya Kothandaraman +3
Traditional video codecs optimized for pixel fidelity collapse at ultra-low bitrates and produce severe artifacts. This failure arises from a fundamental misalignment between pixel…
Zero-Shot Personalized Camera Motion Control for Image-to-Video Synthesis
Pooja Guhan, Divya Kothandaraman, Geonsun Lee +3
Specifying nuanced and compelling camera motion remains a significant hurdle for non-expert creators using generative tools, creating an "expressive gap" where generic text prompts…
3D-free meets 3D priors: Novel View Synthesis from a Single Image with Pretrained Diffusion Guidance
Taewon Kang, Divya Kothandaraman, Dinesh Manocha +1
Recent 3D novel view synthesis (NVS) methods often require extensive 3D data for training, and also typically lack generalization beyond the training distribution. Moreover, they t…
Financial Models in Generative Art: Black-Scholes-Inspired Concept Blending in Text-to-Image Diffusion
Divya Kothandaraman, Ming Lin, Dinesh Manocha
We introduce a novel approach for concept blending in pretrained text-to-image diffusion models, aiming to generate images at the intersection of multiple text prompts. At each tim…