From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
Wei Li, Peijin Jia, Yuan Ma +9
FoMoVLA enhances vision-language-action models by jointly predicting future visual features and tracking sparse 2D points, providing both goal states and motion paths to improve co…
cs.RO2025
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
Titong Jiang, Xuefeng Jiang, Yuan Ma +7
We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in exe…
cs.RO2025
TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving
Xuefeng Jiang, Yuan Ma, Pengxiang Li +7
In recent years, diffusion models have demonstrated remarkable potential across diverse domains, from vision generation to language modeling. Transferring its generative capabiliti…