4 papers
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
Erfei Cui, Wenhai Wang, Zhiqi Li +7
Large language models (LLMs) have opened up new possibilities for intelligent agents, endowing them with human-like thinking and cognitive abilities. In this work, we delve into th…
Demystify Transformers & Convolutions in Modern Image Deep Networks
Xiaowei Hu, Min Shi, Weiyun Wang +9
Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these adva…
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
Yuchen Duan, Weiyun Wang, Zhe Chen +7
Transformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processi…
Parameter-Inverted Image Pyramid Networks
Xizhou Zhu, Xue Yang, Zhaokai Wang +6
Image pyramids are commonly used in modern computer vision tasks to obtain multi-scale features for precise understanding of images. However, image pyramids process multiple resolu…