collaborators

16 papers

cs.RO2026

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

Zhihao Zhu, Hanlin Shang, Mingwang Xu +6

Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high comp…

cs.CV2026

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

Yixuan Lai, Tianjia Shao, Kun Zhou +3

GroundShot is a training-free, model-agnostic framework that improves visual consistency in multi-shot video generation by maintaining an online entity-level visual memory and sche…

cs.CV2026

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture

Hui Li, Fu-Yun Wang, Haoyuan Xia +4

This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one tra…

cs.CV2026

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation

Yuxuan Yao, Yuxuan Chen, Hui Li +6

Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visua…

cs.CV2026

SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation

Weijia Dou, Hui Li, Jiahao Cui +3

Streaming video generation models typically rely on temporal-centric memory, which organizes historical context as raw frames, chunk segments, or unclustered tokens. This organizat…

eess.IV2026

ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality

Feng Ding, Haisheng Fu, Jie Liang +3

We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information for downstream models. We form…