Publications (24)
Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation
Ariel Ephrat, Inbar Mosseri, Oran Lang +5
We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio…
Converse theorems assuming a partial Euler product
David W. Farmer, Kevin Wilson
Associated to a newform is a Dirichlet series with functional equation and Euler product. Hecke showed that if the Dirichlet series has a functional equation…
Investigating Bayesian optimization for expensive-to-evaluate black box functions: Application in fluid dynamics
Mike Diessner, Joseph O'Connor, Andrew Wynn +4
Bayesian optimization provides an effective method to optimize expensive-to-evaluate black box functions. It has been widely applied to problems in many fields, including notably i…
Differentiable Consistency Constraints for Improved Deep Speech Enhancement
Scott Wisdom, John R. Hershey, Kevin Wilson +4
In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement system…
Sequential Multi-Frame Neural Beamforming for Speech Separation and Enhancement
Zhong-Qiu Wang, Hakan Erdogan, Scott Wisdom +5
This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks…
AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech
Brian Patton, Yannis Agiomyrgiannakis, Michael Terry +3
Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opin…
Incomplete resection of the icEEG seizure onset zone is not associated with post-surgical outcomes
Sarah J. Gascoigne, Nathan Evans, Gerard Hall +16
Delineation of seizure onset regions from EEG is important for effective surgical workup. However, it is unknown if their complete resection is required for seizure freedom, or in…
Recomposer: Event-roll-guided generative audio editing
Daniel P. W. Ellis, Eduardo Fonseca, Ron J. Weiss +7
Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their str…
Comparing Methodological Variations in Seizure Onset Localisation Algorithms using intracranial EEG
Sarah J. Gascoigne, Manel Vila-Vidal, Nathan Evans +9
During clinical treatment for epilepsy, the area of the brain thought to be responsible for pathological activity is identified. This identification is typically performed through…
A library of quantitative markers of seizure severity
Sarah J. Gascoigne, Leonard Waldmann, Mariella Panagiotopoulou +16
Purpose: Understanding fluctuations of seizure severity within individuals is important for defining treatment outcomes and response to therapy, as well as developing novel treatme…
Unsupervised Sound Separation Using Mixture Invariant Training
Scott Wisdom, Efthymios Tzinis, Hakan Erdogan +3
In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a…
CNN Architectures for Large-Scale Audio Classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis +10
Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of…
Unsupervised Multi-channel Separation and Adaptation
Cong Han, Kevin Wilson, Scott Wisdom +1
A key challenge in machine learning is to generalize from training data to an application domain of interest. This work generalizes the recently-proposed mixture invariant training…
Conditional Copula models using loss-based Bayesian Additive Regression Trees
Tathagata Basu, Fabrizio Leisen, Cristiano Villa +1
The study of dependence between random variables under external influences is a challenging problem in multivariate analysis. We address this by proposing a novel semi-parametric a…
Inferring Mbh-Mbulge Evolution from the Gravitational Wave Background
Cayenne Matt, Kayhan Gultekin, Luke Kelley +104
We test the impact of an evolving supermassive black hole (SMBH) mass scaling relation (Mbh-Mbulge) on the predictions for the gravitational wave background (GWB). The observed GWB…
Exploring Tradeoffs in Models for Low-latency Speech Enhancement
Kevin Wilson, Michael Chinen, Jeremy Thorpe +5
We explore a variety of neural networks configurations for one- and two-channel spectrogram-mask-based speech enhancement. Our best model improves on previous state-of-the-art perf…
End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings
Soumi Maiti, Hakan Erdogan, Kevin Wilson +3
We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling spe…
Universal Sound Separation
Ilya Kavalerov, Scott Wisdom, Hakan Erdogan +4
Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating…
Private Vertical Federated Inference for Time-Series
Lucas Fenaux, Larris Xie, Aditya Bang +3
Institutions may benefit from collaborative inference on time-series data. In settings where privacy is necessary, multi-party computation (MPC) is a straightforward approach to pr…
Bayesian design and analysis of two-arm cluster randomised trials using assurance: extension to binary outcomes and comparison of MCMC and INLA
Abdullah Aloufi, Kevin Wilson, Nina Wilson +2
The paper considers two different designs; a two-arm superiority cluster randomised controlled trial (RCT) with a continuous outcome, and a twoarm superiority cluster RCT with a bi…
Distance-Based Sound Separation
Katharine Patterson, Kevin Wilson, Scott Wisdom +1
We propose the novel task of distance-based sound separation, where sounds are separated based only on their distance from a single microphone. In the context of assisted listening…
VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition
Quan Wang, Ignacio Lopez Moreno, Mert Saglam +8
We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speec…
VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking
Quan Wang, Hannah Muckenhirn, Kevin Wilson +7
In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We ac…
AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies
Sourish Chaudhuri, Joseph Roth, Daniel P. W. Ellis +8
Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio-…