11 papers · 1 filter
Learning Explicit Physical Parameter Control and Benchmarking for Video Generation
Yanxun Li, Hao Wen, Bingze Song +7
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Cu…
Vidu S1: A Real-Time Interactive Video Generation Model
Jintao Zhang, Kai Jiang, Jintao Chen +24
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment throug…
AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation
Rui Qian, Chuanhang Deng, Qiang Huang +6
Reasoning segmentation requires models to ground complex, implicit textual queries into precise pixel-level masks. Existing approaches rely on a single segmentation token $\texttt{…
Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
Jintao Chen, Chengyu Bai, Junjun Hu +2
Autoregressive video synthesis offers a promising pathway for infinite-horizon generation but is fundamentally hindered by three intertwined challenges: semantic forgetting from co…
AstraNav-World: World Model for Foresight Control and Consistency
Jintao Chen, Junjun Hu, Haochen Bai +11
Embodied navigation in open, dynamic environments demands accurate foresight of how the world will evolve and how actions will unfold over time. We propose AstraNav-World, an end-t…
ConceptWeaver: Weaving Disentangled Concepts with Flow
Jintao Chen, Aiming Hao, Xiaoqing Chen +6
Pre-trained flow-based models excel at synthesizing complex scenes yet lack a direct mechanism for disentangling and customizing their underlying concepts from one-shot real-world…