42 citations · 94 across the 11 of their papers we have counts for
7 papers · 1 filter
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Hao Li, Changyao Tian, Jie Shao +8
The remarkable success of Large Language Models (LLMs) has extended to the multimodal domain, achieving outstanding performance in image understanding and generation. Recent effort…
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
Erfei Cui, Wenhai Wang, Zhiqi Li +7
Large language models (LLMs) have opened up new possibilities for intelligent agents, endowing them with human-like thinking and cognitive abilities. In this work, we delve into th…
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
Changyao Tian, Chenxin Tao, Jifeng Dai +7
Image recognition and generation have long been developed independently of each other. With the recent trend towards general-purpose representation learning, the development of gen…
Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information
Weijie Su, Xizhou Zhu, Chenxin Tao +7
To effectively exploit the potential of large-scale models, various pre-training strategies supported by massive data from different sources are proposed, including supervised pre-…
Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
Hao Li, Jinguo Zhu, Xiaohu Jiang +8
Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to elimi…
Demystify Transformers & Convolutions in Modern Image Deep Networks
Xiaowei Hu, Min Shi, Weiyun Wang +9
Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these adva…