3 citations · 4 across the 3 of their papers we have counts for
16 papers · 1 filter
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
Frank Palma Gomez, Ramon Sanabria, Yun-hsuan Sung +3
Large language models (LLMs) are trained on text-only data that go far beyond the languages with paired speech and text data. At the same time, Dual Encoder (DE) based retrieval sy…
Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
Alexandra Saliba, Yuanchao Li, Ramon Sanabria +1
The efficacy of self-supervised speech models has been validated, yet the optimal utilization of their representations remains challenging across diverse tasks. In this study, we d…
Acoustic Word Embeddings for Untranscribed Target Languages with Continued Pretraining and Learned Pooling
Ramon Sanabria, Ondrej Klejch, Hao Tang +1
Acoustic word embeddings are typically created by training a pooling function using pairs of word-like units. For unsupervised systems, these are mined using k-nearest neighbor (KN…
The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR
Ramon Sanabria, Nikolay Bogoychev, Nina Markl +3
English is the most widely spoken language in the world, used daily by millions of people as a first or second language in many different contexts. As a result, there are many vari…
On the Difficulty of Segmenting Words with Attention
Ramon Sanabria, Hao Tang, Sharon Goldwater
Word segmentation, the problem of finding word boundaries in speech, is of interest for a range of tasks. Previous papers have suggested that for sequence-to-sequence models traine…
Talk, Don't Write: A Study of Direct Speech-Based Image Retrieval
Ramon Sanabria, Austin Waters, Jason Baldridge
Speech-based image retrieval has been studied as a proxy for joint representation learning, usually without emphasis on retrieval itself. As such, it is unclear how well speech-bas…