7 papers · 1 filter
DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving
Zebin Xing, Yupeng Zheng, Qiang Chen +10
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across perception, language, and p…
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
Linbo Wang, Yupeng Zheng, Qiang Chen +13
We introduce Latent-WAM, an efficient end-to-end autonomous driving framework that achieves strong trajectory planning through spatially-aware and dynamics-informed latent world re…
Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
Yichen Zhang, Da Peng, Zonghao Guo +19
A recent cutting-edge topic in multimodal modeling is to unify visual comprehension and generation within a single model. However, the two tasks demand mismatched decoding regimes…
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
Shichu Sun, Yichen Zhang, Haolin Song +6
Visual encoding followed by token condensing has become the standard architectural paradigm in multi-modal large language models (MLLMs). Many recent MLLMs increasingly favor globa…
ContentV: Efficient Training of Video Generation Models with Limited Compute
Wenfeng Lin, Renjie Chen, Boyuan Liu +10
Recent advances in video generation demand increasingly efficient training recipes to mitigate escalating computational costs. In this report, we present ContentV, an 8B-parameter…
Towards Self-Improvement of Diffusion Models via Group Preference Optimization
Renjie Chen, Wenfeng Lin, Yichen Zhang +5
Aligning text-to-image (T2I) diffusion models with Direct Preference Optimization (DPO) has shown notable improvements in generation quality. However, applying DPO to T2I faces two…