Showing eess.ASShow all
2 papers · 1 filter
eess.AS2025
Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning
Manh Luong, Khai Nguyen, Dinh Phung +2
Audio captioning systems face a fundamental challenge: teacher-forcing training creates exposure bias that leads to caption degeneration during inference. While contrastive methods…
eess.AS2021
Many-to-Many Voice Conversion based Feature Disentanglement using Variational Autoencoder
Manh Luong, Viet Anh Tran
Voice conversion is a challenging task which transforms the voice characteristics of a source speaker to a target speaker without changing linguistic content. Recently, there have…