activity
20212026
most citedSemantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding

5 citations · 13 across the 24 of their papers we have counts for

collaborators
Showing eess.ASShow all

8 papers · 1 filter

eess.AS2024

CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR

Wei Zhou, Junteng Jia, Leda Sari +2

CTC compressor can be an effective approach to integrate audio encoders to decoder-only models, which has gained growing interest for different speech applications. In this work, w…

eess.AS2024

Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech

Wonjune Kang, Junteng Jia, Chunyang Wu +8

This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system w…

eess.AS2024

M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses

Yufeng Yang, Desh Raj, Ju Lin +8

The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. Howeve…

eess.AS2024

Faster Speech-LLaMA Inference with Multi-token Prediction

Desh Raj, Gil Keren, Junteng Jia +2

Large language models (LLMs) have become proficient at solving a wide variety of tasks, including those involving multi-modal inputs. In particular, instantiating an LLM (such as L…

eess.AS2024

Effective internal language model training and fusion for factorized transducer model

Jinxi Guo, Niko Moritz, Yingyi Ma +6

The internal language model (ILM) of the neural transducer has been widely studied. In most prior work, it is mainly used for estimating the ILM score and is subsequently subtracte…

eess.AS2023

End-to-End Speech Recognition Contextualization with Large Language Models

Egor Lakomkin, Chunyang Wu, Yassir Fathullah +3

In recent years, Large Language Models (LLMs) have garnered significant attention from the research community due to their exceptional performance and generalization capabilities.…