works on

From the 1 of 26 linked papers with an AI index.

activity
20242026
collaborators

26 papers

cs.AI2026

AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach

Zixuan Jiang, Binghao Qiang, Jiaying Chi +3

The paper introduces AgenticASR, an architecture that continuously refines speech recognition output to remove disfluencies and preserve speaker intent during live audio streams.

eess.AS2026

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

Yuxiang Zhao, Yichi Zhang, Yanjie An +10

Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have…

eess.AS2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang +36

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…

cs.LG2026

Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations

Wentao Hu, Yanbo Zhai, Xiaohui Hu +6

Sparse Mixture-of-Experts (MoE) models have achieved remarkable scalability, yet they remain vulnerable to hallucinations, particularly when processing long-tail knowledge. We iden…

eess.AS2026

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

Qixi Zheng, Yuxiang Zhao, Tianrui Wang +7

Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have…

eess.AS2026

Anonymization, Not Elimination: Utility-Preserved Speech Anonymization

Yunchong Xiao, Yuxiang Zhao, Ziyang Ma +4

The growing reliance on large-scale speech data has made privacy protection a critical concern. However, existing anonymization approaches often degrade data utility, for example b…