5 papers
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
Madhuri Shanbhogue, Zhe Li, Shanfeng Zhang +86
We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage t…
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
Jaedong Hwang, Kumar Tanmay, Seok-Jin Lee +5
Large Language Models (LLMs) have achieved strong performance in domains like mathematics, factual question answering, and code generation, yet their ability to reason on these tas…
Sign Language: Towards Sign Understanding for Robot Autonomy
Ayush Agrawal, Joel Loo, Nicky Zimmerman +1
Navigational signs are common aids for human wayfinding and scene understanding, but are underutilized by robots. We argue that they benefit robot navigation and scene understandin…
SignLoc: Robust Localization using Navigation Signs and Public Maps
Nicky Zimmerman, Joel Loo, Ayush Agrawal +1
Navigation signs and maps, such as floor plans and street maps, are widely available and serve as ubiquitous aids for way-finding in human environments. Yet, they are rarely used b…
Language Models' Factuality Depends on the Language of Inquiry
Tushar Aggarwal, Kumar Tanmay, Ayush Agrawal +3
Multilingual language models (LMs) are expected to recall factual knowledge consistently across languages, yet they often fail to transfer knowledge between languages even when the…