papers

Publications (11)

cs.CL2016

Discriminative Acoustic Word Embeddings: Recurrent Neural Network-Based Approaches

Shane Settle, Karen Livescu

Acoustic word embeddings --- fixed-dimensional vector representations of variable-length spoken word segments --- have begun to be considered for tasks such as speech recognition a…

cs.CL2017

Query-by-Example Search with Discriminative Neural Acoustic Word Embeddings

Shane Settle, Keith Levin, Herman Kamper +1

Query-by-example search often uses dynamic time warping (DTW) for comparing queries and proposed matching segments. Recent work has shown that comparing speech segments by represen…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

cs.CL2023

Neural approaches to spoken content embedding

Shane Settle

Comparing spoken segments is a central operation to speech processing. Traditional approaches in this area have favored frame-level dynamic programming algorithms, such as dynamic…

cs.CL2020

Acoustic span embeddings for multilingual query-by-example search

Yushi Hu, Shane Settle, Karen Livescu

Query-by-example (QbE) speech search is the task of matching spoken queries to utterances within a search collection. In low- or zero-resource settings, QbE search is often address…

cs.CL2019

Acoustically Grounded Word Embeddings for Improved Acoustics-to-Word Speech Recognition

Shane Settle, Kartik Audhkhasi, Karen Livescu +1

Direct acoustics-to-word (A2W) systems for end-to-end automatic speech recognition are simpler to train, and more efficient to decode with, than sub-word systems. However, A2W syst…