4 papers
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
Madhuri Shanbhogue, Zhe Li, Shanfeng Zhang +86
We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage t…
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
Ke Wu, Ehsan Variani, Tom Bagby +2
We present MedASR, an open-source 105M-parameter model engineered for high-accuracy medical dictation. Prioritizing a "small, fast, and accurate" design, MedASR addresses 3 core pi…
Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)
Cyril Allauzen, Tom Bagby, Georg Heigold +2
The Massive Sound Embedding Benchmark (MSEB) has emerged as a standard for evaluating the functional breadth of audio models. While initial baselines focused on specialized encoder…
Massive Sound Embedding Benchmark (MSEB)
Georg Heigold, Ehsan Variani, Tom Bagby +4
Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcri…