40 citations · 61 across the 7 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2023
BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic Data
Xuenan Xu, Zhiling Zhang, Zelin Zhou +4
Compared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data. Most existing metho…
cs.SD2021★ 3 cited
Enriching Ontology with Temporal Commonsense for Low-Resource Audio Tagging
Zhiling Zhang, Zelin Zhou, Haifeng Tang +3
Audio tagging aims at predicting sound events occurred in a recording. Traditional models require enormous laborious annotations, otherwise performance degeneration will be the nor…