206 citations · 382 across the 17 of their papers we have counts for
6 papers · 1 filter
Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
Jizhong Liu, Gang Li, Junbo Zhang +5
Automated audio captioning (AAC) is an audio-to-text task to describe audio contents in natural language. Recently, the advancements in large language models (LLMs), with improveme…
Bridging Language Gaps in Audio-Text Retrieval
Zhiyong Yan, Heinrich Dinkel, Yongqing Wang +4
Audio-text retrieval is a challenging task, requiring the search for an audio clip or a text caption within a database. The predominant focus of existing research on English descri…
Scaling up masked audio encoder learning for general audio classification
Heinrich Dinkel, Zhiyong Yan, Yongqing Wang +3
Despite progress in audio classification, a generalization gap remains between speech and other sound domains, such as environmental sounds and music. Models trained for speech tas…
CED: Consistent ensemble distillation for audio tagging
Heinrich Dinkel, Yongqing Wang, Zhiyong Yan +2
Augmentation and knowledge distillation (KD) are well-established techniques employed in audio classification tasks, aimed at enhancing performance and reducing model sizes on the…
Understanding temporally weakly supervised training: A case study for keyword spotting
Heinrich Dinkel, Weiji Zhuang, Zhiyong Yan +3
The currently most prominent algorithm to train keyword spotting (KWS) models with deep neural networks (DNNs) requires strong supervision i.e., precise knowledge of the spoken key…
Unified Keyword Spotting and Audio Tagging on Mobile Devices with Transformers
Heinrich Dinkel, Yongqing Wang, Zhiyong Yan +2
Keyword spotting (KWS) is a core human-machine-interaction front-end task for most modern intelligent assistants. Recently, a unified (UniKW-AT) framework has been proposed that ad…