papers

Publications (49)

cs.CL2021

Cascaded Multilingual Audio-Visual Learning from Videos

Andrew Rouditchenko, Angie Boggust, David Harwath +8

In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to vis…

cs.CV2021

Routing with Self-Attention for Multimodal Capsule Networks

Kevin Duarte, Brian Chen, Nina Shvetsova +7

The task of multimodal learning has seen a growing interest recently as it allows for training neural architectures based on different modalities such as vision, text, and audio. O…

cs.CL2019

English Broadcast News Speech Recognition by Humans and Machines

Samuel Thomas, Masayuki Suzuki, Yinghui Huang +8

With recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversa…

cs.CV2021

Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Brian Chen, Andrew Rouditchenko, Kevin Duarte +10

Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data…

stat.CO2020

Learning Hamiltonian Monte Carlo in R

Samuel Thomas, Wanzhu Tu

Hamiltonian Monte Carlo (HMC) is a powerful tool for Bayesian computation. In comparison with the traditional Metropolis-Hastings algorithm, HMC offers greater computational effici…

eess.AS2026

Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts

George Saon, Samuel Thomas, Takashi Fukuda +3

We propose self-speculative decoding for speech-aware LLMs by using the CTC encoder as a draft model to accelerate auto-regressive (AR) inference and improve ASR accuracy. Our thre…

eess.AS2026

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

Haolong Zheng, Yuanzhuo Hu, Xinyu Liang +7

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and l…

cs.CL2016

Invariant Representations for Noisy Speech Recognition

Dmitriy Serdyuk, Kartik Audhkhasi, Philémon Brakel +3

Modern automatic speech recognition (ASR) systems need to be robust under acoustic variability arising from environmental, speaker, channel, and recording conditions. Ensuring such…

cs.SD2025

A Non-autoregressive Model for Joint STT and TTS

Vishal Sunder, Brian Kingsbury, George Saon +5

In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimoda…

eess.AS2025

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

George Saon, Avihu Dekel, Alexander Brooks +21

Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…

cs.CL2023

Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages

Andrew Rouditchenko, Sameer Khurana, Samuel Thomas +6

Recent models such as XLS-R and Whisper have made multilingual speech technologies more accessible by pre-training on audio from around 100 spoken languages each. However, there ar…

cs.CL2018

SimplerVoice: A Key Message & Visual Description Generator System for Illiteracy

Minh N. B. Nguyen, Samuel Thomas, Anne E. Gattiker +2

We introduce SimplerVoice: a key message and visual description generator system to help low-literate adults navigate the information-dense world with confidence, on their own. Sim…

cs.CL2022

Integrating Text Inputs For Training and Adapting RNN Transducer ASR Models

Samuel Thomas, Brian Kingsbury, George Saon +1

Compared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be independently adapted to a new domain, recent end-to-en…

hep-lat2017

Structure of the Nucleon and its Excitations

Waseem Kamleh, Derek Leinweber, Zhan-wei Liu +4

The structure of the ground state nucleon and its finite-volume excitations are examined from three different perspectives. Using new techniques to extract the relativistic compone…

cs.CL2018

A Recorded Debating Dataset

Shachar Mirkin, Michal Jacovi, Tamar Lavee +6

This paper describes an English audio and textual dataset of debating speeches, a unique resource for the growing research field of computational argumentation and debating technol…

eess.AS2022

Extending RNN-T-based speech recognition systems with emotion and language classification

Zvi Kons, Hagai Aronowitz, Edmilson Morais +4

Speech transcription, emotion recognition, and language identification are usually considered to be three different tasks. Each one requires a different model with a different arch…

eess.AS2026

NLE: Non-autoregressive LLM-based ASR by Transcript Editing

Avihu Dekel, Samuel Thomas, Takashi Fukada +1

While autoregressive (AR) LLM-based ASR systems achieve strong accuracy, their sequential decoding limits parallelism and incurs high latency. We propose NLE, a non-autoregressive…

cs.CL2022

Improving End-to-End Models for Set Prediction in Spoken Language Understanding

Hong-Kwang J. Kuo, Zoltan Tuske, Samuel Thomas +2

The goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts…

cs.CL2021

RNN Transducer Models For Spoken Language Understanding

Samuel Thomas, Hong-Kwang J. Kuo, George Saon +5

We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding(SLU). These end-to-end (E2E) models are constructed in thr…

cs.MM2025

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

Edson Araujo, Andrew Rouditchenko, Yuan Gong +7

Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations t…

cs.CL2020

End-to-End Spoken Language Understanding Without Full Transcripts

Hong-Kwang J. Kuo, Zoltán Tüske, Samuel Thomas +7

An essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develo…

eess.AS2025

Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?

Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo +4

We propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new St…

cs.CV2024

What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions

Brian Chen, Nina Shvetsova, Andrew Rouditchenko +6

Spatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually traine…

cs.CV2018

Understanding Unequal Gender Classification Accuracy from Face Images

Vidya Muthukumar, Tejaswini Pedapati, Nalini Ratha +7

Recent work shows unequal performance of commercial face classification services in the gender classification task across intersectional groups defined by skin type and gender. Acc…

cs.PL2021

A Compiler Infrastructure for Accelerator Generators

Rachit Nigam, Samuel Thomas, Zhijing Li +1

We present Calyx, a new intermediate language (IL) for compiling high-level programs into hardware designs. Calyx combines a hardware-like structural language with a software-like…

cs.CV2026

TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning

Soumya Shamarao Jahagirdar, Edson Araujo, Anna Kukleva +7

Recent video reasoning models have shown strong results on temporal and multimodal understanding, yet they depend on large-scale supervised data and multi-stage training pipelines,…

cs.DC2020

Using Skip Graphs for Increased NUMA Locality

Samuel Thomas, Ana Hayne, Jonad Pulaj +1

We present a data partitioning technique performed over skip graphs that promotes significant quantitative and qualitative improvements on NUMA locality in concurrent data structur…

cs.CL2021

Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible Inputs

Sujeong Cha, Wangrui Hou, Hyun Jung +5

A major focus of recent research in spoken language understanding (SLU) has been on the end-to-end approach where a single model can predict intents directly from speech inputs wit…

cs.CL2018

Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition

Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg +3

The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significa…

eess.AS2025

Towards Audio Token Compression in Large Audio Language Models

Saurabhchand Bhati, Samuel Thomas, Hilde Kuehne +2

Large Audio Language Models (LALMs) demonstrate impressive performance across diverse tasks, ranging from speech recognition to general audio understanding. However, their scalabil…

cs.CL2023

C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval

Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova +7

Multilingual text-video retrieval methods have improved significantly in recent years, but the performance for other languages lags behind English. We propose a Cross-Lingual Cross…

cs.CV2022

Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

Nina Shvetsova, Brian Chen, Andrew Rouditchenko +6

Multi-modal learning from video data has seen increased attention recently as it allows to train semantically meaningful embeddings without human annotation enabling tasks like zer…

cs.CL2022

Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding

Vishal Sunder, Samuel Thomas, Hong-Kwang J. Kuo +3

Dialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in tex…

cs.CL2021

Integrating Dialog History into End-to-End Spoken Language Understanding Systems

Jatin Ganhotra, Samuel Thomas, Hong-Kwang J. Kuo +4

End-to-end spoken language understanding (SLU) systems that process human-human or human-computer interactions are often context independent and process each turn of a conversation…

cs.CL2020

End-to-end spoken language understanding using transformer networks and self-supervised pre-trained features

Edmilson Morais, Hong-Kwang J. Kuo, Samuel Thomas +2

Transformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in t…

cs.CL2017

English Conversational Telephone Speech Recognition by Humans and Machines

George Saon, Gakuto Kurata, Tom Sercu +9

One of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major sp…

eess.AS2024

Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Andrew Rouditchenko, Yuan Gong, Samuel Thomas +4

Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is…

cs.DS2024

Lower Bounds for Approximate (& Exact) k-Disjoint-Shortest-Paths

Rajesh Chitnis, Samuel Thomas, Anthony Wirth

Given a graph and a set of pairs, the -vertex-disjoint-paths (resp. -edge-disjoint-paths) problem asks t…

cs.PL2020

Predictable Accelerator Design with Time-Sensitive Affine Types

Rachit Nigam, Sachille Atapattu, Samuel Thomas +6

Field-programmable gate arrays (FPGAs) provide an opportunity to co-design applications with hardware accelerators, yet they remain difficult to program. High-level synthesis (HLS)…

cs.CV2026

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers

Edson Araujo, Saurabhchand Bhati, M. Jehanzeb Mirza +5

Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over…

cs.CL2022

Towards Reducing the Need for Speech Training Data To Build Spoken Language Understanding Systems

Samuel Thomas, Hong-Kwang J. Kuo, Brian Kingsbury +1

The lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly proces…

cs.CL2022

A new data augmentation method for intent classification enhancement and its application on spoken conversation datasets

Zvi Kons, Aharon Satt, Hong-Kwang Kuo +4

Intent classifiers are vital to the successful operation of virtual agent systems. This is especially so in voice activated systems where the data can be noisy with many ambiguous…

cs.CL2020

Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems

Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas +5

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is…

cs.CV2023

FisHook -- An Optimized Approach to Marine Specie Classification using MobileNetV2

Kohav Dey, Krishna Bajaj, K S Ramalakshmi +2

Marine ecosystems are vital for the planet's health, but human activities such as climate change, pollution, and overfishing pose a constant threat to marine species. Accurate clas…

eess.AS2026

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions

Xulin Fan, Vishal Sunder, Samuel Thomas +3

Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer…

eess.AS2025

mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition

Andrew Rouditchenko, Samuel Thomas, Hilde Kuehne +2

Audio-Visual Speech Recognition (AVSR) combines lip-based video with audio and can improve performance in noise, but most methods are trained only on English data. One limitation i…

stat.AP2024

Predictive Modeling of Lower-Level English Club Soccer Using Crowd-Sourced Player Valuations

Josh Brown, Yutong Bu, Zachary Cheesman +5

In this research, we examine the capabilities of different mathematical models to accurately predict various levels of the English football pyramid. Existing work has largely focus…

cs.CV2021

AVLnet: Learning Audio-Visual Language Representations from Instructional Videos

Andrew Rouditchenko, Angie Boggust, David Harwath +11

Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognitio…

cs.CL2022

Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent Systems

Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas +2

Recent advances in End-to-End (E2E) Spoken Language Understanding (SLU) have been primarily due to effective pretraining of speech representations. One such pretraining paradigm is…