collaborators

6 papers

cs.AI2025

PortAgent: LLM-driven Vehicle Dispatching Agent for Port Terminals

Jia Hu, Junqi Li, Weimeng Lin +3

Vehicle Dispatching Systems (VDSs) are critical to the operational efficiency of Automated Container Terminals (ACTs). However, their widespread commercialization is hindered due t…

cs.RO2025

The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning

Titong Jiang, Xuefeng Jiang, Yuan Ma +7

We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in exe…

cs.CV2025

World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

Yupeng Zheng, Pengxuan Yang, Zebin Xing +8

End-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information.…

cs.CV2025

DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models

Yuhan Hao, Zhengning Li, Lei Sun +7

Vision-Language-Action (VLA) models have advanced autonomous driving, but existing benchmarks still lack scenario diversity, reliable action-level annotation, and evaluation protoc…

cs.RO2025

TransDiffuser: Diverse Trajectory Generation with Decorrelated Multi-modal Representation for End-to-end Autonomous Driving

Xuefeng Jiang, Yuan Ma, Pengxiang Li +7

In recent years, diffusion models have demonstrated remarkable potential across diverse domains, from vision generation to language modeling. Transferring its generative capabiliti…

cs.CV2025

TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference

Junshan Hu, Jialiang Mao, Zhikang Liu +3

Conventional Vision-Language Models(VLMs) typically utilize a fixed number of vision tokens, regardless of task complexity. This one-size-fits-all strategy introduces notable ineff…