Publications (50)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
Keqi Deng, Jinxi Guo, Yingyi Ma +4
Biased Self-supervised learning for ASR
Florian L. Kreyssig, Yangyang Shi, Jinxi Guo +3
A Distributed Optimisation Framework Combining Natural Gradient with Hessian-Free for Discriminative Sequence Training
Adnan Haider, Chao Zhang, Florian L. Kreyssig +1
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
Keqi Deng, Guangzhi Sun, Philip C. Woodland
Integrating Source-channel and Attention-based Sequence-to-sequence Models for Speech Recognition
Qiujia Li, Chao Zhang, Philip C. Woodland
Self-supervised representations in speech-based depression detection
Wen Wu, Chao Zhang, Philip C. Woodland
End-to-end Spoken Language Understanding with Tree-constrained Pointer Generator
Guangzhi Sun, Chao Zhang, Philip C. Woodland
Discriminative Neural Clustering for Speaker Diarisation
Qiujia Li, Florian L. Kreyssig, Chao Zhang +1
SOT Triggered Neural Clustering for Speaker Attributed ASR
Xianrui Zheng, Guangzhi Sun, Chao Zhang +1
FastInject: Injecting Unpaired Text Data into CTC-based ASR training
Keqi Deng, Philip C. Woodland
Distribution-based Emotion Recognition in Conversation
Wen Wu, Chao Zhang, Philip C. Woodland
Sequence Training of DNN Acoustic Models With Natural Gradient
Adnan Haider, Philip C. Woodland
Estimating the Uncertainty in Emotion Class Labels with Utterance-Specific Dirichlet Priors
Wen Wu, Chao Zhang, Xixin Wu +1
Improved Large-margin Softmax Loss for Speaker Diarisation
Yassir Fathullah, Chao Zhang, Philip C. Woodland
Cross-Lingual Interleaving for Speech Language Models
Adel Moumen, Guangzhi Sun, Philip C. Woodland
Knowledge-Aware Audio-Grounded Generative Slot Filling for Limited Annotated Data
Guangzhi Sun, Chao Zhang, Ivan VuliÄ +2
Speech-based Slot Filling using Large Language Models
Guangzhi Sun, Shutong Feng, Dongcheng Jiang +3
DNCASR: End-to-End Training for Speaker-Attributed ASR
Xianrui Zheng, Chao Zhang, Philip C. Woodland
Decoupled Structure for Improved Adaptability of End-to-End Models
Keqi Deng, Philip C. Woodland
Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription
Xianrui Zheng, Chao Zhang, Philip C. Woodland
Label-Synchronous Neural Transducer for End-to-End ASR
Keqi Deng, Philip C. Woodland
Adaptable End-to-End ASR Models using Replaceable Internal LMs and Residual Softmax
Keqi Deng, Philip C. Woodland
Emotion recognition by fusing time synchronous and time asynchronous representations
Wen Wu, Chao Zhang, Philip C. Woodland
Confidence Estimation for Automatic Detection of Depression and Alzheimer's Disease Based on Clinical Interviews
Wen Wu, Chao Zhang, Philip C. Woodland
Can Contextual Biasing Remain Effective with Whisper and GPT-2?
Guangzhi Sun, Xianrui Zheng, Chao Zhang +1
Knowledge Distillation for Neural Transducers from Large Self-Supervised Pre-trained Models
Xiaoyu Yang, Qiujia Li, Philip C. Woodland
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
Keqi Deng, Wenxi Chen, Xie Chen +1
1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem
Mingjie Chen, Hezhao Zhang, Yuanchao Li +11
Knowledge Distillation from Multiple Foundation Models for End-to-End Speech Recognition
Xiaoyu Yang, Qiujia Li, Chao Zhang +1
Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning
Mohan Li, Rama Doddipatla, Philip C. Woodland
Tree-constrained Pointer Generator for End-to-end Contextual Speech Recognition
Guangzhi Sun, Chao Zhang, Philip C. Woodland
Multi-head Temporal Latent Attention
Keqi Deng, Philip C. Woodland
Self-Supervised Learning-Based Source Separation for Meeting Data
Yuang Li, Xianrui Zheng, Philip C. Woodland
Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
Wen Wu, Chao Zhang, Philip C. Woodland
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
Keqi Deng, Philip C. Woodland
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
Keqi Deng, Philip C. Woodland
Confidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognition
Qiujia Li, David Qiu, Yu Zhang +5
Combining Frame-Synchronous and Label-Synchronous Systems for Speech Recognition
Qiujia Li, Chao Zhang, Philip C. Woodland
Tree-constrained Pointer Generator with Graph Neural Network Encodings for Contextual Speech Recognition
Guangzhi Sun, Chao Zhang, Philip C. Woodland
Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for Conversations
Wen Wu, Chao Zhang, Philip C. Woodland
Spectral Clustering-aware Learning of Embeddings for Speaker Diarisation
Evonne P. C. Lee, Guangzhi Sun, Chao Zhang +1
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
Mengqi Wang, Zhan Liu, Zengrui Jin +3
Handling Ambiguity in Emotion: From Out-of-Domain Detection to Distribution Estimation
Wen Wu, Bo Li, Chao Zhang +5
CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
Guangzhi Sun, Xiao Zhan, Shutong Feng +2
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
Nineli Lashkarashvili, Wen Wu, Guangzhi Sun +1
Adapting GPT, GPT-2 and BERT Language Models for Speech Recognition
Xianrui Zheng, Chao Zhang, Philip C. Woodland
Residual Energy-Based Models for End-to-End Speech Recognition
Qiujia Li, Yu Zhang, Bo Li +2
Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition
Qiujia Li, Yu Zhang, David Qiu +3
It HAS to be Subjective: Human Annotator Simulation via Zero-shot Density Estimation
Wen Wu, Wenlin Chen, Chao Zhang +1
Cosine-Distance Virtual Adversarial Training for Semi-Supervised Speaker-Discriminative Acoustic Embeddings
Florian L. Kreyssig, Philip C. Woodland