Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
Hao Shi, Yusuke Fujita, Tomoya Mizumoto +3
Prompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic spee…
cs.CL2024
Self-Supervised Learning for Multi-Channel Neural Transducer
Atsushi Kojima
Self-supervised learning, such as with the wav2vec 2.0 framework significantly improves the accuracy of end-to-end automatic speech recognition (ASR). Wav2vec 2.0 has been applied…
cs.CL2024
Intermediate direct preference optimization
Atsushi Kojima
We propose the intermediate direct preference optimization (DPO) method to calculate the DPO loss at selected intermediate layers as an auxiliary loss for finetuning large language…