◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

K. Yu

39 papers hereh-index 7025.2k citations945 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author6
  • middle author20
  • last author12

Across the 38 of 39 papers where every author was matched, so the position is known.

fields
  • cs.CV10
  • cs.SD7
  • cs.LG5
  • eess.AS4
  • cs.RO3
  • cs.CL2
same name
  • K. Yu — 39 papers, h 25
  • K. Yu — 13 papers, h 30
  • K. Yu — 4 papers, h 5
  • K. Yu — 4 papers, h 58
  • K. Yu — 3 papers, h 7
  • K. Yu — 3 papers, h 1

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20082026
most citedEDVR: Video Restoration with Enhanced Deformable Convolutional Networks

62 citations · 190 across the 31 of their papers we have counts for

collaborators
Showing eess.ASShow all

4 papers · 1 filter

eess.AS2023

Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS

Yifan Yang, Feiyu Shen, Chenpeng Du +4

Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer…

eess.AS2023

VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching

Yiwei Guo, Chenpeng Du, Ziyang Ma +2

Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms th…

eess.AS2023

Improving Audio Caption Fluency with Automatic Error Correction

Hanxue Zhang, Zeyu Xie, Xuenan Xu +2

Automated audio captioning (AAC) is an important cross-modality translation task, aiming at generating descriptions for audio clips. However, captions generated by previous AAC mod…

eess.AS2023

DiffVoice: Text-to-Speech with Latent Diffusion

Zhijun Liu, Yiwei Guo, Kai Yu

In this work, we present DiffVoice, a novel text-to-speech model based on latent diffusion. We propose to first encode speech signals into a phoneme-rate latent representation with…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.