papers

Publications (24)

cs.SD2018

Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

Ariel Ephrat, Inbar Mosseri, Oran Lang +5

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio…

math.NT2005

Converse theorems assuming a partial Euler product

David W. Farmer, Kevin Wilson

Associated to a newform is a Dirichlet series with functional equation and Euler product. Hecke showed that if the Dirichlet series has a functional equation…

cs.LG2022

Investigating Bayesian optimization for expensive-to-evaluate black box functions: Application in fluid dynamics

Mike Diessner, Joseph O'Connor, Andrew Wynn +4

Bayesian optimization provides an effective method to optimize expensive-to-evaluate black box functions. It has been widely applied to problems in many fields, including notably i…

cs.SD2018

Differentiable Consistency Constraints for Improved Deep Speech Enhancement

Scott Wisdom, John R. Hershey, Kevin Wilson +4

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement system…

cs.SD2020

Sequential Multi-Frame Neural Beamforming for Speech Separation and Enhancement

Zhong-Qiu Wang, Hakan Erdogan, Scott Wisdom +5

This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks…

cs.CL2016

AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech

Brian Patton, Yannis Agiomyrgiannakis, Michael Terry +3

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opin…

q-bio.NC2023

Incomplete resection of the icEEG seizure onset zone is not associated with post-surgical outcomes

Sarah J. Gascoigne, Nathan Evans, Gerard Hall +16

Delineation of seizure onset regions from EEG is important for effective surgical workup. However, it is unknown if their complete resection is required for seizure freedom, or in…

cs.SD2025

Recomposer: Event-roll-guided generative audio editing

Daniel P. W. Ellis, Eduardo Fonseca, Ron J. Weiss +7

Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their str…

q-bio.NC2024

Comparing Methodological Variations in Seizure Onset Localisation Algorithms using intracranial EEG

Sarah J. Gascoigne, Manel Vila-Vidal, Nathan Evans +9

During clinical treatment for epilepsy, the area of the brain thought to be responsible for pathological activity is identified. This identification is typically performed through…

q-bio.NC2023

A library of quantitative markers of seizure severity

Sarah J. Gascoigne, Leonard Waldmann, Mariella Panagiotopoulou +16

Purpose: Understanding fluctuations of seizure severity within individuals is important for defining treatment outcomes and response to therapy, as well as developing novel treatme…

eess.AS2020

Unsupervised Sound Separation Using Mixture Invariant Training

Scott Wisdom, Efthymios Tzinis, Hakan Erdogan +3

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a…

cs.SD2017

CNN Architectures for Large-Scale Audio Classification

Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis +10

Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of…

cs.SD2024

Unsupervised Multi-channel Separation and Adaptation

Cong Han, Kevin Wilson, Scott Wisdom +1

A key challenge in machine learning is to generalize from training data to an application domain of interest. This work generalizes the recently-proposed mixture invariant training…

stat.ME2026

Conditional Copula models using loss-based Bayesian Additive Regression Trees

Tathagata Basu, Fabrizio Leisen, Cristiano Villa +1

The study of dependence between random variables under external influences is a challenging problem in multivariate analysis. We address this by proposing a novel semi-parametric a…

astro-ph.HE2025

Inferring Mbh-Mbulge Evolution from the Gravitational Wave Background

Cayenne Matt, Kayhan Gultekin, Luke Kelley +104

We test the impact of an evolving supermassive black hole (SMBH) mass scaling relation (Mbh-Mbulge) on the predictions for the gravitational wave background (GWB). The observed GWB…

cs.SD2018

Exploring Tradeoffs in Models for Low-latency Speech Enhancement

Kevin Wilson, Michael Chinen, Jeremy Thorpe +5

We explore a variety of neural networks configurations for one- and two-channel spectrogram-mask-based speech enhancement. Our best model improves on previous state-of-the-art perf…

cs.SD2021

End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

Soumi Maiti, Hakan Erdogan, Kevin Wilson +3

We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling spe…

cs.SD2019

Universal Sound Separation

Ilya Kavalerov, Scott Wisdom, Hakan Erdogan +4

Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating…

cs.LG2026

Private Vertical Federated Inference for Time-Series

Lucas Fenaux, Larris Xie, Aditya Bang +3

Institutions may benefit from collaborative inference on time-series data. In settings where privacy is necessary, multi-party computation (MPC) is a straightforward approach to pr…

stat.ME2025

Bayesian design and analysis of two-arm cluster randomised trials using assurance: extension to binary outcomes and comparison of MCMC and INLA

Abdullah Aloufi, Kevin Wilson, Nina Wilson +2

The paper considers two different designs; a two-arm superiority cluster randomised controlled trial (RCT) with a continuous outcome, and a twoarm superiority cluster RCT with a bi…

cs.SD2022

Distance-Based Sound Separation

Katharine Patterson, Kevin Wilson, Scott Wisdom +1

We propose the novel task of distance-based sound separation, where sounds are separated based only on their distance from a single microphone. In the context of assisted listening…

eess.AS2020

VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition

Quan Wang, Ignacio Lopez Moreno, Mert Saglam +8

We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speec…

eess.AS2019

VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Quan Wang, Hannah Muckenhirn, Kevin Wilson +7

In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We ac…

cs.SD2018

AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies

Sourish Chaudhuri, Joseph Roth, Daniel P. W. Ellis +8

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio-…