Publications (49)
Cascaded Multilingual Audio-Visual Learning from Videos
Andrew Rouditchenko, Angie Boggust, David Harwath +8
In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to vis…
Routing with Self-Attention for Multimodal Capsule Networks
Kevin Duarte, Brian Chen, Nina Shvetsova +7
The task of multimodal learning has seen a growing interest recently as it allows for training neural architectures based on different modalities such as vision, text, and audio. O…
English Broadcast News Speech Recognition by Humans and Machines
Samuel Thomas, Masayuki Suzuki, Yinghui Huang +8
With recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversa…
Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos
Brian Chen, Andrew Rouditchenko, Kevin Duarte +10
Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data…
Learning Hamiltonian Monte Carlo in R
Samuel Thomas, Wanzhu Tu
Hamiltonian Monte Carlo (HMC) is a powerful tool for Bayesian computation. In comparison with the traditional Metropolis-Hastings algorithm, HMC offers greater computational effici…
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
George Saon, Samuel Thomas, Takashi Fukuda +3
We propose self-speculative decoding for speech-aware LLMs by using the CTC encoder as a draft model to accelerate auto-regressive (AR) inference and improve ASR accuracy. Our thre…
CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling
Haolong Zheng, Yuanzhuo Hu, Xinyu Liang +7
CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and l…
Invariant Representations for Noisy Speech Recognition
Dmitriy Serdyuk, Kartik Audhkhasi, Philémon Brakel +3
Modern automatic speech recognition (ASR) systems need to be robust under acoustic variability arising from environmental, speaker, channel, and recording conditions. Ensuring such…
A Non-autoregressive Model for Joint STT and TTS
Vishal Sunder, Brian Kingsbury, George Saon +5
In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimoda…
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
George Saon, Avihu Dekel, Alexander Brooks +21
Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…
Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages
Andrew Rouditchenko, Sameer Khurana, Samuel Thomas +6
Recent models such as XLS-R and Whisper have made multilingual speech technologies more accessible by pre-training on audio from around 100 spoken languages each. However, there ar…
SimplerVoice: A Key Message & Visual Description Generator System for Illiteracy
Minh N. B. Nguyen, Samuel Thomas, Anne E. Gattiker +2
We introduce SimplerVoice: a key message and visual description generator system to help low-literate adults navigate the information-dense world with confidence, on their own. Sim…
Integrating Text Inputs For Training and Adapting RNN Transducer ASR Models
Samuel Thomas, Brian Kingsbury, George Saon +1
Compared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be independently adapted to a new domain, recent end-to-en…
Structure of the Nucleon and its Excitations
Waseem Kamleh, Derek Leinweber, Zhan-wei Liu +4
The structure of the ground state nucleon and its finite-volume excitations are examined from three different perspectives. Using new techniques to extract the relativistic compone…
A Recorded Debating Dataset
Shachar Mirkin, Michal Jacovi, Tamar Lavee +6
This paper describes an English audio and textual dataset of debating speeches, a unique resource for the growing research field of computational argumentation and debating technol…
Extending RNN-T-based speech recognition systems with emotion and language classification
Zvi Kons, Hagai Aronowitz, Edmilson Morais +4
Speech transcription, emotion recognition, and language identification are usually considered to be three different tasks. Each one requires a different model with a different arch…
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
Avihu Dekel, Samuel Thomas, Takashi Fukada +1
While autoregressive (AR) LLM-based ASR systems achieve strong accuracy, their sequential decoding limits parallelism and incurs high latency. We propose NLE, a non-autoregressive…
Improving End-to-End Models for Set Prediction in Spoken Language Understanding
Hong-Kwang J. Kuo, Zoltan Tuske, Samuel Thomas +2
The goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts…
RNN Transducer Models For Spoken Language Understanding
Samuel Thomas, Hong-Kwang J. Kuo, George Saon +5
We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding(SLU). These end-to-end (E2E) models are constructed in thr…
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
Edson Araujo, Andrew Rouditchenko, Yuan Gong +7
Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations t…
End-to-End Spoken Language Understanding Without Full Transcripts
Hong-Kwang J. Kuo, Zoltán Tüske, Samuel Thomas +7
An essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develo…
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo +4
We propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new St…
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
Brian Chen, Nina Shvetsova, Andrew Rouditchenko +6
Spatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually traine…
Understanding Unequal Gender Classification Accuracy from Face Images
Vidya Muthukumar, Tejaswini Pedapati, Nalini Ratha +7
Recent work shows unequal performance of commercial face classification services in the gender classification task across intersectional groups defined by skin type and gender. Acc…
A Compiler Infrastructure for Accelerator Generators
Rachit Nigam, Samuel Thomas, Zhijing Li +1
We present Calyx, a new intermediate language (IL) for compiling high-level programs into hardware designs. Calyx combines a hardware-like structural language with a software-like…
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
Soumya Shamarao Jahagirdar, Edson Araujo, Anna Kukleva +7
Recent video reasoning models have shown strong results on temporal and multimodal understanding, yet they depend on large-scale supervised data and multi-stage training pipelines,…
Using Skip Graphs for Increased NUMA Locality
Samuel Thomas, Ana Hayne, Jonad Pulaj +1
We present a data partitioning technique performed over skip graphs that promotes significant quantitative and qualitative improvements on NUMA locality in concurrent data structur…
Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible Inputs
Sujeong Cha, Wangrui Hou, Hyun Jung +5
A major focus of recent research in spoken language understanding (SLU) has been on the end-to-end approach where a single model can predict intents directly from speech inputs wit…
Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg +3
The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significa…
Towards Audio Token Compression in Large Audio Language Models
Saurabhchand Bhati, Samuel Thomas, Hilde Kuehne +2
Large Audio Language Models (LALMs) demonstrate impressive performance across diverse tasks, ranging from speech recognition to general audio understanding. However, their scalabil…
C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval
Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova +7
Multilingual text-video retrieval methods have improved significantly in recent years, but the performance for other languages lags behind English. We propose a Cross-Lingual Cross…
Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval
Nina Shvetsova, Brian Chen, Andrew Rouditchenko +6
Multi-modal learning from video data has seen increased attention recently as it allows to train semantically meaningful embeddings without human annotation enabling tasks like zer…
Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding
Vishal Sunder, Samuel Thomas, Hong-Kwang J. Kuo +3
Dialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in tex…
Integrating Dialog History into End-to-End Spoken Language Understanding Systems
Jatin Ganhotra, Samuel Thomas, Hong-Kwang J. Kuo +4
End-to-end spoken language understanding (SLU) systems that process human-human or human-computer interactions are often context independent and process each turn of a conversation…
End-to-end spoken language understanding using transformer networks and self-supervised pre-trained features
Edmilson Morais, Hong-Kwang J. Kuo, Samuel Thomas +2
Transformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in t…
English Conversational Telephone Speech Recognition by Humans and Machines
George Saon, Gakuto Kurata, Tom Sercu +9
One of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major sp…
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
Andrew Rouditchenko, Yuan Gong, Samuel Thomas +4
Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is…
Lower Bounds for Approximate (& Exact) k-Disjoint-Shortest-Paths
Rajesh Chitnis, Samuel Thomas, Anthony Wirth
Given a graph and a set of pairs, the -vertex-disjoint-paths (resp. -edge-disjoint-paths) problem asks t…
Predictable Accelerator Design with Time-Sensitive Affine Types
Rachit Nigam, Sachille Atapattu, Samuel Thomas +6
Field-programmable gate arrays (FPGAs) provide an opportunity to co-design applications with hardware accelerators, yet they remain difficult to program. High-level synthesis (HLS)…
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
Edson Araujo, Saurabhchand Bhati, M. Jehanzeb Mirza +5
Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over…
Towards Reducing the Need for Speech Training Data To Build Spoken Language Understanding Systems
Samuel Thomas, Hong-Kwang J. Kuo, Brian Kingsbury +1
The lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly proces…
A new data augmentation method for intent classification enhancement and its application on spoken conversation datasets
Zvi Kons, Aharon Satt, Hong-Kwang Kuo +4
Intent classifiers are vital to the successful operation of virtual agent systems. This is especially so in voice activated systems where the data can be noisy with many ambiguous…
Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems
Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas +5
Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is…
FisHook -- An Optimized Approach to Marine Specie Classification using MobileNetV2
Kohav Dey, Krishna Bajaj, K S Ramalakshmi +2
Marine ecosystems are vital for the planet's health, but human activities such as climate change, pollution, and overfishing pose a constant threat to marine species. Accurate clas…
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
Xulin Fan, Vishal Sunder, Samuel Thomas +3
Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer…
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
Andrew Rouditchenko, Samuel Thomas, Hilde Kuehne +2
Audio-Visual Speech Recognition (AVSR) combines lip-based video with audio and can improve performance in noise, but most methods are trained only on English data. One limitation i…
Predictive Modeling of Lower-Level English Club Soccer Using Crowd-Sourced Player Valuations
Josh Brown, Yutong Bu, Zachary Cheesman +5
In this research, we examine the capabilities of different mathematical models to accurately predict various levels of the English football pyramid. Existing work has largely focus…
AVLnet: Learning Audio-Visual Language Representations from Instructional Videos
Andrew Rouditchenko, Angie Boggust, David Harwath +11
Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognitio…
Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent Systems
Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas +2
Recent advances in End-to-End (E2E) Spoken Language Understanding (SLU) have been primarily due to effective pretraining of speech representations. One such pretraining paradigm is…