activity
20182023
most citedHierarchical Cross-Modal Talking Face Generationwith Dynamic Pixel-Wise Loss

27 citations · 33 across the 12 of their papers we have counts for

collaborators

16 papers

eess.AS2023

Mitigating Cross-Database Differences for Learning Unified HRTF Representation

Yutong Wen, You Zhang, Zhiyao Duan

Individualized head-related transfer functions (HRTFs) are crucial for accurate sound positioning in virtual auditory displays. As the acoustic measurement of HRTFs is resource-int…

eess.AS2023

SingNet: A Real-time Singing Voice Beat and Downbeat Tracking System

Mojtaba Heydari, Ju-Chiang Wang, Zhiyao Duan

Singing voice beat and downbeat tracking posses several applications in automatic music production, analysis and manipulation. Among them, some require real-time processing, such a…

cs.SD2023

Phase perturbation improves channel robustness for speech spoofing countermeasures

Yongyi Zang, You Zhang, Zhiyao Duan

In this paper, we aim to address the problem of channel robustness in speech countermeasure (CM) systems, which are used to distinguish synthetic speech from human natural speech.…

eess.AS2022

SAMO: Speaker Attractor Multi-Center One-Class Learning for Voice Anti-Spoofing

Siwen Ding, You Zhang, Zhiyao Duan

Voice anti-spoofing systems are crucial auxiliaries for automatic speaker verification (ASV) systems. A major challenge is caused by unseen attacks empowered by advanced speech syn…

eess.AS20222 cited

A Data-Driven Methodology for Considering Feasibility and Pairwise Likelihood in Deep Learning Based Guitar Tablature Transcription Systems

Frank Cwitkowitz, Jonathan Driedger, Zhiyao Duan

Guitar tablature transcription is an important but understudied problem within the field of music information retrieval. Traditional signal processing approaches offer only limited…

eess.AS2021

A study of the robustness of raw waveform based speaker embeddings under mismatched conditions

Ge Zhu, Frank Cwitkowitz, Zhiyao Duan

In this paper, we conduct a cross-dataset study on parametric and non-parametric raw-waveform based speaker embeddings through speaker verification experiments. In general, we obse…