3 papers
cs.RO2025
WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving
Pengxuan Yang, Ben Lu, Zhongpu Xia +7
Latent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. Howe…
cs.RO2025
TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving
Xuefeng Jiang, Yuan Ma, Pengxiang Li +7
In recent years, diffusion models have demonstrated remarkable potential across diverse domains, from vision generation to language modeling. Transferring its generative capabiliti…
cs.CV2025
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
Junshan Hu, Jialiang Mao, Zhikang Liu +3
Conventional Vision-Language Models(VLMs) typically utilize a fixed number of vision tokens, regardless of task complexity. This one-size-fits-all strategy introduces notable ineff…