activity
20162024
most citedDeep Spatio-Temporal Residual Networks for Citywide Crowd Flows Prediction

206 citations · 382 across the 17 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2024

Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding

Jizhong Liu, Gang Li, Junbo Zhang +5

Automated audio captioning (AAC) is an audio-to-text task to describe audio contents in natural language. Recently, the advancements in large language models (LLMs), with improveme…

cs.SD2024

Bridging Language Gaps in Audio-Text Retrieval

Zhiyong Yan, Heinrich Dinkel, Yongqing Wang +4

Audio-text retrieval is a challenging task, requiring the search for an audio clip or a text caption within a database. The predominant focus of existing research on English descri…

cs.SD2024

Scaling up masked audio encoder learning for general audio classification

Heinrich Dinkel, Zhiyong Yan, Yongqing Wang +3

Despite progress in audio classification, a generalization gap remains between speech and other sound domains, such as environmental sounds and music. Models trained for speech tas…

cs.SD2023

CED: Consistent ensemble distillation for audio tagging

Heinrich Dinkel, Yongqing Wang, Zhiyong Yan +2

Augmentation and knowledge distillation (KD) are well-established techniques employed in audio classification tasks, aimed at enhancing performance and reducing model sizes on the…

cs.SD2023

Understanding temporally weakly supervised training: A case study for keyword spotting

Heinrich Dinkel, Weiji Zhuang, Zhiyong Yan +3

The currently most prominent algorithm to train keyword spotting (KWS) models with deep neural networks (DNNs) requires strong supervision i.e., precise knowledge of the spoken key…

cs.SD2023

Unified Keyword Spotting and Audio Tagging on Mobile Devices with Transformers

Heinrich Dinkel, Yongqing Wang, Zhiyong Yan +2

Keyword spotting (KWS) is a core human-machine-interaction front-end task for most modern intelligent assistants. Recently, a unified (UniKW-AT) framework has been proposed that ad…