From the 4 of 45 linked papers with an AI index.
45 papers
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
Dongxu Ge, Shansong Liu, Cheng Gong +3
As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted incre…
CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
Yuyang Huang, Yabo Chen, Wenrui Dai +6
CineWeaver introduces a training-free method that modifies pretrained video diffusion models to generate long, multi-shot cinematic videos with fine-grained reference control and c…
ShotPlan: Cinematic Video Generation with Learnable Planning Token
Su Guo, Guangce Liu, Haosen Yang +7
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective mult…
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Rui Li, Yuanzhi Liang, Ziqi Ni +3
The paper proposes TaRoS, a framework that redesigns reward signals for video generation using GRPO to avoid reward hacking and saturation, improving visual fidelity, motion cohere…
SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
Ruoyu Wang, Jialun Liu, Huayang Huang +5
The paper introduces Self-Imagination Fine-Tuning (SIFT), a method that trains video diffusion models on their own generated videos to improve physical realism and disentangle moti…
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Yuanzhi Liang, Xufeng Zhan, Haibin Huang +2
The paper outlines a roadmap for building open‑world physical intelligence by integrating World Action Models with an "embodied brain" architecture that unifies multimodal context,…