From the 1 of 6 linked papers with an AI index.
6 papers
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Xu Guo, Zhengxuan Wei, Xinghui Li +11
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs.…
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting
Jiwen Liu, Shujuan Li, Xiaohan Li +5
The paper introduces TARS, a 3D‑free video re‑shooting framework that uses text‑driven semantic viewpoint specifications and self‑supervised training to control camera motion and p…
Vera: Identity-Faithful Human Subject-to-Video Generation
Yulong Xu, Xinyue Liu, Shujuan Li +6
Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for…
ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation
Zijie Meng, Jiwen Liu, Yufei Liu +5
Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression s…
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
Yuan Wang, Ouxiang Li, Yulong Xu +8
Recent advances in generative video models are increasingly driven by post-training and test-time scaling, both of which critically depend on the quality of video reward models (RM…
FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems
Borui Liao, Yulong Xu, Jiao Ou +4
Full-Duplex Speech Dialogue Systems (Full-Duplex SDS) have significantly enhanced the naturalness of human-machine interaction by enabling real-time bidirectional communication. Ho…