activity
20242026
most citedDiscrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Efficient One-to-Many Translation with Joint Multi-Stream Diffusion

Yiwen Guan, Jacob Whitehill

One-to-many machine translation (MT) is computationally expensive for autoregressive (AR) systems, which suffer from linear latency scaling with both sequence length and the number…

cs.CL2026

Learning to Translate from Soft to Hard LLM Prompts

Pitipat Kongsomjit, Suryansh Goyal, Jacob Whitehill

Soft prompting, also known as continuous prompting, is a parameter-efficient method for tuning LLMs to specific tasks. Like other machine learning techniques, its parameters encode…

cs.CL2025

Interactive In-Meeting Speaker Correction with Human Feedback

Xinlu He, Yiwen Guan, Badrivishal Paurana +3

Most automatic speech processing systems operate in ``open loop'' mode without user feedback about who said what, yet human-in-the-loop workflows can potentially enable higher accu…

cs.CL2025

Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation

Yiwen Guan, Jacob Whitehill

Multilingual translation suffers from computational redundancy, especially when translating into multiple languages simultaneously. In addition, translation quality can suffer for…

cs.CL2025

Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context

Viet Anh Trinh, Xinlu He, Jacob Whitehill

Classroom speech and lectures often contain named entities (NEs) such as names of people and special terminology. While automatic speech recognition (ASR) systems have achieved rem…

cs.CL2025

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio

Xinlu He, Jacob Whitehill

Monaural multi-speaker automatic speech recognition (ASR) remains challenging due to data scarcity and the intrinsic difficulty of recognizing and attributing words to individual s…