most citedMambaBEV: An EV-based 3D detection model with Mamba2

1 citations · 1 across the 1 of their papers we have counts for

collaborators

8 papers

cs.CV20261 cited

MambaBEV: An EV-based 3D detection model with Mamba2

Zihan You, Ni Wang, Hao Wang +2

Accurate 3D object detection in autonomous driving relies on Bird's Eye View (BEV) perception and effective temporal fusion. However, existing fusion strategies based on convolutio…

cs.CV2026

Kelix Technical Report

Boyang Ding, Chenglong Chu, Dunju Zang +28

Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which u…

cs.CV2025

HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices

HyperAI Team, Yuchen Liu, Kaiyang Han +26

Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy direc…

cs.CV2025

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

Hao Wang, Eiki Murata, Lingfang Zhang +11

Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet…

cs.CV2025

EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence

Ding Zou, Feifan Wang, Mengyu Ge +17

The realization of Artificial General Intelligence (AGI) necessitates Embodied AI agents capable of robust spatial perception, effective task planning, and adaptive execution in ph…

cs.RO2025

Igniting VLMs toward the Embodied Space

Andy Zhai, Brae Liu, Bruno Fang +17

While foundation models show remarkable progress in language and vision, existing vision-language models (VLMs) still have limited spatial and embodiment understanding. Transferrin…