From the 1 of 7 linked papers with an AI index.
7 papers
TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models
Taewon Kang, Matthias Zwicker
The paper introduces Temporal Prior Decoupling (TPD), a training‑free method that restores suppressed late‑segment information during diffusion sampling for text‑to‑video models, i…
Text-Conditioned Background Generation for Editable Multi-Layer Documents
Taewon Kang, Joseph K J, Chris Tensmeyer +4
We present a framework for document-centric background generation with multi-page editing and thematic continuity. To ensure text regions remain readable, we employ a latent maskin…
Character-Centered Dialogue Generation from Scene-Level Prompts
Taewon Kang, Ming C. Lin
Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- character-driven dialogue and speech…
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
Taewon Kang, Divya Kothandaraman, Ming C. Lin
Generating coherent long-form video sequences from discrete text prompts remains challenging due to difficulties in maintaining temporal coherence, semantic consistency, and scene-…
DCR: Counterfactual Attractor Guidance for Rare Compositional Generation
Taewon Kang, Matthias Zwicker
Diffusion models generate realistic visual content, yet often fail to produce rare but plausible compositions. When prompted with combinations that are valid but underrepresented i…
NEGATE: Constrained Semantic Guidance for Linguistic Negation in Text-to-Video Diffusion
Taewon Kang, Ming C. Lin
Negation is a fundamental linguistic operator, yet it remains inadequately modeled in diffusion-based generative systems. In this work, we present a formal treatment of linguistic…