Publications (9)
The NUS-HLT System for ICASSP2024 ICMC-ASR Grand Challenge
Meng Ge, Yizhou Peng, Yidi Jiang +6
This paper summarizes our team's efforts in both tracks of the ICMC-ASR Challenge for in-car multi-channel automatic speech recognition. Our submitted systems for ICMC-ASR Challeng…
Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation
Wupeng Wang, Zexu Pan, Jingru Lin +2
Speech separation seeks to isolate individual speech signals from a multi-talk speech mixture. Despite much progress, a system well-trained on synthetic data often experiences perf…
Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture Speech
Jingru Lin, Meng Ge, Wupeng Wang +2
Self-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-label…
Self-Supervised Acoustic Word Embedding Learning via Correspondence Transformer Encoder
Jingru Lin, Xianghu Yue, Junyi Ao +1
Acoustic word embeddings (AWEs) aims to map a variable-length speech segment into a fixed-dimensional representation. High-quality AWEs should be invariant to variations, such as d…
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
Jingru Lin, Chen Zhang, Tianrui Wang +1
Due to recent advancements in Large Audio-Language Models (LALMs) that demonstrate remarkable performance across a range of sound-, speech- and music-related tasks, there is a grow…
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
Jingru Lin, Chen Zhang, Stephen Y. Liu +1
Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallucinations-by dynamically retriev…
Reinforcement Learning Foundations for Deep Research Systems: A Survey
Wenjun Li, Zhi Chen, Jingru Lin +8
Deep research systems, agentic AI that solve complex, multi-step tasks by coordinating reasoning, search across the open web and user files, and tool use, are moving toward hierarc…
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
Jingru Lin, Meng Ge, Junyi Ao +2
It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on singl…
Interpolating Speaker Identities in Embedding Space for Data Expansion
Tianchi Liu, Ruijie Tao, Qiongqiong Wang +5
The success of deep learning-based speaker verification systems is largely attributed to access to large-scale and diverse speaker identity data. However, collecting data from more…