collaborators

20 papers

cs.DC2026

Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures

Haibo HU, Lianming Huang, Qiao Li +2

Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they intr…

cs.CV2026

Mode-as-Sequence: Translating Multimodal Motion Prediction into Unified Sequential Mode Modeling

Zikang Zhou, Haibo Hu, Xinhong Chen +5

Multimodal motion forecasting is inherently under-supervised: each training scene provides only one realized future, yet multiple plausible futures exist. This sparse supervision o…

cs.CL2026

Retrieval-Augmented Generation for Natural Language Processing: A Survey

Shangyu Wu, Ying Xiong, Yufei Cui +8

Large language models (LLMs) have achieved strong empirical performance in various fields, benefiting from their huge amount of parameters that store knowledge. However, LLMs still…

cs.CL2026

RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference

Lianming Huang, Shangyu Wu, Yufei Cui +6

Deploying large language model inference remains challenging due to their high computational overhead. Early exit optimizes model inference by adaptively reducing the number of inf…

cs.CL2026

ReFilter: Improving Robustness of Retrieval-Augmented Generation via Gated Filter

Yixin Chen, Ying Xiong, Shangyu Wu +3

Retrieval-augmented generation (RAG) has become a dominant paradigm for grounding large language models (LLMs) with external evidence in knowledge-intensive question answering. A c…

cs.CV2025

DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving

Haibo HU, Lianming Huang, Nan Guan +1

Vision-Language Action (VLA) models unify perception, reasoning, and trajectory generation for autonomous driving, but suffer from significant inference latency due to deep transfo…