4 papers
Learning the joint distribution of two sequences using little or no paired data
Soroosh Mariooryad, Matt Shannon, Siyuan Ma +5
We present a noisy channel generative model of two sequences, for example text and speech, which enables uncovering the association between the two modalities when limited paired d…
On exponential convergence of SGD in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, Siyuan Ma
Large over-parametrized models learned via stochastic gradient descent (SGD) methods have become a key element in modern machine learning. Although SGD methods are very effective i…
Kernel Machines Beat Deep Neural Networks on Mask-based Single-channel Speech Enhancement
Like Hui, Siyuan Ma, Mikhail Belkin
We apply a fast kernel method for mask-based single-channel speech enhancement. Specifically, our method solves a kernel regression problem associated to a non-smooth kernel functi…
Kernel machines that adapt to GPUs for effective large batch training
Siyuan Ma, Mikhail Belkin
Modern machine learning models are typically trained using Stochastic Gradient Descent (SGD) on massively parallel computing resources such as GPUs. Increasing mini-batch size is a…