18 papers
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
Jiayi Luo, Hanxin Zhu, Chen Gao +5
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…
AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment
Ziyao Huang, Shunkai Li, Juan Cao +7
Recent advances in video diffusion models have spurred interest in human-object interaction (HOI) video generation, which demands fine-grained control over interaction logic beyond…
Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment
Zixiang Zhou, Zhentao Yu, Yifeng Ma +9
Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity consistency and model comple…
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Sen Liang, Cong Wang, Zhentao Yu +8
Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge…
HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation
Cong Wang, Zhentao Yu, Hongmei Wang +7
Current identity-consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi-view reference input offers a n…
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
Cong Wan, Zeyu Guo, Zijian Cai +6
Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics…