6 citations · 6 across the 2 of their papers we have counts for
5 papers · 1 filter
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt
Yanfeng Shi, Pengfei Cai, Jun Liu +5
Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenge…
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
Pengfei Cai, Yan Song, Qing Gu +3
Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts hav…
An Efficient Transfer Learning Method Based on Adapter with Local Attributes for Speech Emotion Recognition
Haoyu Song, Ian McLoughlin, Qing Gu +2
Existing speech emotion recognition (SER) methods commonly suffer from the lack of high-quality large-scale corpus, partly due to the complex, psychological nature of emotion which…
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
Pengfei Cai, Yan Song, Nan Jiang +2
A significant challenge in sound event detection (SED) is the effective utilization of unlabeled data, given the limited availability of labeled data due to high annotation costs.…
USTC-KXDIGIT System Description for ASVspoof5 Challenge
Yihao Chen, Haochen Wu, Nan Jiang +13
This paper describes the USTC-KXDIGIT system submitted to the ASVspoof5 Challenge for Track 1 (speech deepfake detection) and Track 2 (spoofing-robust automatic speaker verificatio…