2 papers
cs.SD2026
Uni-ASR: Unified LLM-Based Architecture for Non-Streaming and Streaming Automatic Speech Recognition
Yinfeng Xia, Jian Tang, Junfeng Hou +2
Although the deep integration of the Automatic Speech Recognition (ASR) system with Large Language Models (LLMs) has significantly improved accuracy, the deployment of such systems…
eess.AS2018
A Capsule based Approach for Polyphonic Sound Event Detection
Yaming Liu, Jian Tang, Yan Song +1
Polyphonic sound event detection (polyphonic SED) is an interesting but challenging task due to the concurrence of multiple sound events. Recently, SED methods based on convolution…