◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Y. Xiao

4 papers hereh-index 64k citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3
  • last author1

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL3
  • cs.CV1
same name
  • Y. Xiao — 107 papers, h 26
  • Y. Xiao — 15 papers, h 22
  • Y. Xiao — 6 papers, h 17
  • Y. Xiao — 2 papers
  • Y. Xiao — 1 paper, h 1
  • Y. Xiao — 1 paper, h 6

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20172019
most citedUncovering Latent Style Factors for Expressive Speech Synthesis

44 citations · 44 across the 1 of their papers we have counts for

collaborators

4 papers

cs.CV2019

Towards Unconstrained End-to-End Text Spotting

Siyang Qin, Alessandro Bissacco, Michalis Raptis +2

We propose an end-to-end trainable network that can simultaneously detect and recognize text of arbitrary shape, making substantial progress on the open problem of reading scene te…

cs.CL2018

Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron

RJ Skerry-Ryan, Eric Battenberg, Ying Xiao +6

We present an extension to the Tacotron speech synthesis architecture that learns a latent embedding space of prosody, derived from a reference acoustic representation containing t…

cs.CL2018

Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

Yuxuan Wang, Daisy Stanton, Yu Zhang +7

In this work, we propose "global style tokens" (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system. The emb…

cs.CL2017★ 44 cited

Uncovering Latent Style Factors for Expressive Speech Synthesis

Yuxuan Wang, RJ Skerry-Ryan, Ying Xiao +5

Prosodic modeling is a core problem in speech synthesis. The key challenge is producing desirable prosody from textual input containing only phonetic information. In this prelimina…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.