activity
20182022
most citedSynt++: Utilizing Imperfect Synthetic Data to Improve Speech Recognition

4 citations · 5 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CV2022

I see what you hear: a vision-inspired method to localize words

Mohammad Samragh, Arnav Kundu, Ting-Yao Hu +5

This paper explores the possibility of using visual object detection techniques for word localization in speech data. Object detection has been thoroughly studied in the contempora…

eess.AS20214 cited

Synt++: Utilizing Imperfect Synthetic Data to Improve Speech Recognition

Ting-Yao Hu, Mohammadreza Armandpour, Ashish Shrivastava +3

With recent advances in speech synthesis, synthetic data is becoming a viable alternative to real data for training speech recognition models. However, machine learning with synthe…

cs.CV2021

Subspace Representation Learning for Few-shot Image Classification

Ting-Yao Hu, Zhi-Qi Cheng, Alexander G. Hauptmann

In this paper, we propose a subspace representation learning (SRL) framework to tackle few-shot image classification tasks. It exploits a subspace in local CNN feature space to rep…

cs.CV2021

Pose Guided Person Image Generation with Hidden p-Norm Regression

Ting-Yao Hu, Alexander G. Hauptmann

In this paper, we propose a novel approach to solve the pose guided person image generation task. We assume that the relation between pose and appearance information can be describ…

cs.LG20201 cited

SapAugment: Learning A Sample Adaptive Policy for Data Augmentation

Ting-Yao Hu, Ashish Shrivastava, Jen-Hao Rick Chang +5

Data augmentation methods usually apply the same augmentation (or a mix of them) to all the training samples. For example, to perturb data with noise, the noise is sampled from a N…

eess.AS2020

Unsupervised Style and Content Separation by Minimizing Mutual Information for Speech Synthesis

Ting-Yao Hu, Ashish Shrivastava, Oncel Tuzel +1

We present a method to generate speech from input text and a style vector that is extracted from a reference speech signal in an unsupervised manner, i.e., no style annotation, suc…