48 citations · 75 across the 4 of their papers we have counts for
11 papers
Can Diffusion Models Disentangle? A Theoretical Perspective
Liming Wang, Muhammad Jehanzeb Mirza, Yishu Gong +6
This paper presents a novel theoretical framework for understanding how diffusion models can learn disentangled representations. Within this framework, we establish identifiability…
Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment
Yuan Gong, Ziyi Chen, Iek-Heng Chu +2
Automatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency,…
CMKD: CNN/Transformer-Based Cross-Model Knowledge Distillation for Audio Classification
Yuan Gong, Sameer Khurana, Andrew Rouditchenko +1
Audio classification is an active research area with a wide range of applications. Over the past decade, convolutional neural networks (CNNs) have been the de-facto standard buildi…
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, James Glass
In the past decade, convolutional neural networks (CNNs) have been widely adopted as the main building block for end-to-end audio classification models, which aim to learn a direct…
Detecting Replay Attacks Using Multi-Channel Audio: A Neural Network-Based Method
Yuan Gong, Jian Yang, Christian Poellabauer
With the rapidly growing number of security-sensitive systems that use voice as the primary input, it becomes increasingly important to address these systems' potential vulnerabili…
Second-order Non-local Attention Networks for Person Re-identification
Bryan, Xia, Yuan Gong +2
Recent efforts have shown promising results for person re-identification by designing part-based architectures to allow a neural network to learn discriminative representations fro…