3 papers
cs.SD2026
Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features
Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai +1
Recent Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription rely on Grapheme-to-Phoneme (G2P) labels, but the phoneme labels are not necessarily phonetically faithful.…
eess.AS2026
Progressive Alignment Objectives for Aligner-Encoder based ASR
Jaeyoung Lee, Masato Mimura, Takafumi Moriya
Aligner-Encoders are recently proposed seq2seq end-to-end ASR models that replace decoder attention by predicting the uth token directly from the u-th encoder position, so the enco…
eess.AS2026
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
Jaeyoung Lee, Masato Mimura
We present a decoder-only Conformer for automatic speech recognition (ASR) that processes speech and text in a single stack without external speech encoders or pretrained large lan…