16 papers
WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA
Zhihao Zhu, Hanlin Shang, Mingwang Xu +6
Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high comp…
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
Yixuan Lai, Tianjia Shao, Kun Zhou +3
GroundShot is a training-free, model-agnostic framework that improves visual consistency in multi-shot video generation by maintaining an online entity-level visual memory and sche…
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
Hui Li, Fu-Yun Wang, Haoyuan Xia +4
This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one tra…
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation
Yuxuan Yao, Yuxuan Chen, Hui Li +6
Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visua…
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
Weijia Dou, Hui Li, Jiahao Cui +3
Streaming video generation models typically rely on temporal-centric memory, which organizes historical context as raw frames, chunk segments, or unclustered tokens. This organizat…
ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality
Feng Ding, Haisheng Fu, Jie Liang +3
We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information for downstream models. We form…