3 papers
cs.CL2026
Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition
Nathan Roll, Pranav Bhalerao, Martijn Bartelds +7
In speech language modeling, two architectures dominate the frontier: the Transformer and the Conformer. However, it remains unknown whether their comparable performance stems from…
cs.CL2025
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
Nathan Roll, Calbert Graham, Yuka Tatsumi +3
Human listeners readily adjust to unfamiliar speakers and language varieties through exposure, but do these adaptation benefits extend to state-of-the-art spoken language models? W…
cs.SD2025
Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning
Qian Yang, Calbert Graham
Voice conversion (VC) modifies voice characteristics while preserving linguistic content. This paper presents the Stepback network, a novel model for converting speaker identity us…