4 papers
Tracing the Flow of Knowledge From Science to Technology Using Deep Learning
Michael E. Rose, Mainak Ghosh, Sebastian Erhardt +3
We develop a language similarity model suitable for working with patents and scientific publications at the same time. In a horse race-style evaluation, we subject eight language (…
AVD2: Accident Video Diffusion for Accident Video Description
Cheng Li, Keyuan Zhou, Tong Liu +5
Traffic accidents present complex challenges for autonomous driving, often featuring unpredictable scenarios that hinder accurate system interpretation and responses. Nonetheless,…
World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving
Mingliang Zhai, Cheng Li, Zengyuan Guo +7
The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. Howeve…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…