2 papers
cs.CV2024
Fast-BEV: A Fast and Strong Bird's-Eye View Perception Baseline
Yangguang Li, Bin Huang, Zeren Chen +8
Recently, perception task based on Bird's-Eye View (BEV) representation has drawn more and more attention, and BEV representation is promising as the foundation for next-generation…
cs.CV2024
Emu: Generative Pretraining in Multimodality
Quan Sun, Qiying Yu, Yufeng Cui +7
We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-mo…