2 papers
cs.CL2024
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
cs.CV2024
LocCa: Visual Pretraining with Location-aware Captioners
Bo Wan, Michael Tschannen, Yongqin Xian +7
Image captioning has been shown as an effective pretraining method similar to contrastive pretraining. However, the incorporation of location-aware information into visual pretrain…