5 papers · 1 filter
QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation
Jiaqi Zhao, Xiaobin Hu, Bo Yin +3
KV cache memory has become a major deployment bottleneck for video generation and world models, which motivates low-bit quantization study for efficiency. Existing 2-bit KV cache q…
OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation
Zijie Meng, Yufei Liu, Chengqian Ma +8
Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses resi…
KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing
Mingshu Cai, Miao Zhang, Chenghe Yang +3
In recent years, training-free video generation has progressed remarkably. However, when handling complex textual instructions, existing methods still suffer from semantic ambiguit…
DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer
Junpeng Jiang, Gangyi Hong, Miao Zhang +4
Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generati…
DiVE: DiT-based Video Generation with Enhanced Control
Junpeng Jiang, Gangyi Hong, Lijun Zhou +10
Generating high-fidelity, temporally consistent videos in autonomous driving scenarios faces a significant challenge, e.g. problematic maneuvers in corner cases. Despite recent vid…