4 papers
RouteRelay: Event-Triggered Cross-Layer Route Reuse for Efficient Dynamic Sparse Attention
Bin Li, Sisi Liu, Chenyang Hu +3
Dynamic sparse attention reduces long-context prefill cost by routing each query chunk to a small set of key chunks at every Transformer layer. The sparse attention kernel avoids m…
EdgeFM: Efficient Edge Inference for Vision-Language Models
Mengling Deng, Yuanpeng Chen, Sheng Yang +12
Vision-language models (VLMs) have demonstrated strong applicability in edge industrial applications, yet their deployment remains severely constrained by requirements for determin…
Fast-BEV++: Fast by Algorithm, Deployable by Design
Yuanpeng Chen, Hui Song, Sheng Yang +5
The advancement of vision-only BEV (Bird's-Eye-View) perception is hindered by the fundamental trade-off between perception accuracy and deployment efficiency. We introduce Fast-BE…
Precise Drive with VLM: First Prize Solution for PRCV 2024 Drive LM challenge
Bin Huang, Siyu Wang, Yuanpeng Chen +8
This technical report outlines the methodologies we applied for the PRCV Challenge, focusing on cognition and decision-making in driving scenarios. We employed InternVL-2.0, a pion…