#text-to-video generation
5 papers · 1 filter
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Haodong Li, Tianfei Ren, Xiaoxiao Ma +25
The paper presents VideoCoCo, a system that generates physically consistent videos by having a coding agent produce executable Blender code that defines the scene and its dynamics,…
TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models
Taewon Kang, Matthias Zwicker
The paper introduces Temporal Prior Decoupling (TPD), a training‑free method that restores suppressed late‑segment information during diffusion sampling for text‑to‑video models, i…
Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models
Wenxuan Chen, Wenjie Feng
The paper introduces SIRUS, a training‑free, inference‑time method that suppresses specified concepts in text‑to‑video generators while preserving other content, and proposes a vid…
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
Shihao Zhang, Yunzhi Li, Yuguang Yan +4
The paper introduces Shell-LCC, a method that treats the data manifold of high‑quality video training data as an implicit reward model, providing cheap, dense guidance for text‑to‑…
Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation
Ruize Xia
The paper introduces Text2Sign, a diffusion-based model that generates short, low‑resolution sign‑language video clips from text using a single GPU by leveraging a frozen vision‑la…