6 papers
PortAgent: LLM-driven Vehicle Dispatching Agent for Port Terminals
Jia Hu, Junqi Li, Weimeng Lin +3
Vehicle Dispatching Systems (VDSs) are critical to the operational efficiency of Automated Container Terminals (ACTs). However, their widespread commercialization is hindered due t…
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
Titong Jiang, Xuefeng Jiang, Yuan Ma +7
We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in exe…
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
Yupeng Zheng, Pengxuan Yang, Zebin Xing +8
End-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information.…
DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models
Yuhan Hao, Zhengning Li, Lei Sun +7
Vision-Language-Action (VLA) models have advanced autonomous driving, but existing benchmarks still lack scenario diversity, reliable action-level annotation, and evaluation protoc…
TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving
Xuefeng Jiang, Yuan Ma, Pengxiang Li +7
In recent years, diffusion models have demonstrated remarkable potential across diverse domains, from vision generation to language modeling. Transferring its generative capabiliti…
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
Junshan Hu, Jialiang Mao, Zhikang Liu +3
Conventional Vision-Language Models(VLMs) typically utilize a fixed number of vision tokens, regardless of task complexity. This one-size-fits-all strategy introduces notable ineff…