4 citations · 5 across the 10 of their papers we have counts for
7 papers · 1 filter
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Haoqin Sun, Jingguang Tian, Jiaming Zhou +8
The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the…
Learning Emotion-Invariant Speaker Representations for Speaker Verification
Jingguang Tian, Xinhui Hu, Xinkang Xu
In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such repre…
Discrete Audio Representations for Automated Audio Captioning
Jingguang Tian, Haoqin Sun, Xinhui Hu +1
Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous…
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
Jingguang Tian, Shuaishuai Ye, Shunfei Chen +4
This paper presents our system submission for the In-Car Multi-Channel Automatic Speech Recognition (ICMC-ASR) Challenge, which focuses on speaker diarization and speech recognitio…
Large-Scale Learning on Overlapped Speech Detection: New Benchmark and New General System
Zhaohui Yin, Jingguang Tian, Xinhui Hu +2
Overlapped Speech Detection (OSD) is an important part of speech applications involving analysis of multi-party conversations. However, most of existing OSD systems are trained and…
The Royalflush System for VoxCeleb Speaker Recognition Challenge 2022
Jingguang Tian, Xinhui Hu, Xinkang Xu
In this technical report, we describe the Royalflush submissions for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). Our submissions contain track 1, which is for supe…