4 papers
Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
Ze Feng, Jiang-jiang Liu, Sen Yang +5
The computational expense of redundant vision tokens in Large Vision-Language Models (LVLMs) has led many existing methods to compress them via a vision projector. However, this co…
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
Huan Liu, Lingyu Xiao, Jiangjiang Liu +4
With the rapid advancement of Multimodal Large Language Models (MLLMs), a variety of benchmarks have been introduced to evaluate their capabilities. While most evaluations have foc…
Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving
Lingyu Xiao, Jiang-Jiang Liu, Sen Yang +4
The autoregressive world model exhibits robust generalization capabilities in vectorized scene understanding but encounters difficulties in deriving actions due to insufficient unc…
EasyChauffeur: A Baseline Advancing Simplicity and Efficiency on Waymax
Lingyu Xiao, Jiang-Jiang Liu, Xiaoqing Ye +2
Recent advancements in deep-learning-based driving planners have primarily focused on elaborate network engineering, yielding limited improvements. This paper diverges from convent…