Showing cs.SDShow all
2 papers · 1 filter
cs.SD2024
EDTC: enhance depth of text comprehension in automated audio captioning
Liwen Tan, Yin Cao, Yi Zhou
Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multi-modal domains. Facilitating models in…
cs.SD2023
Balanced SNR-Aware Distillation for Guided Text-to-Audio Generation
Bingzhi Liu, Yin Cao, Haohe Liu +1
Diffusion models have demonstrated promising results in text-to-audio generation tasks. However, their practical usability is hindered by slow sampling speeds, limiting their appli…