3 papers
cs.SD2026
Whisfusion: Parallel ASR Decoding with Masked Diffusion
Taeyoun Kwon, Junhyuk Ahn, Taegeun Yun +7
Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length. A natural a…
cs.CL2025
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
Sieun Hyeon, Kyudan Jung, Jaehee Won +4
In various academic and professional settings, such as mathematics lectures or research presentations, it is often necessary to convey mathematical expressions orally. However, rea…
cs.AI2025
MathReader : Text-to-Speech for Mathematical Documents
Sieun Hyeon, Kyudan Jung, Nam-Joon Kim +2
TTS (Text-to-Speech) document reader from Microsoft, Adobe, Apple, and OpenAI have been serviced worldwide. They provide relatively good TTS results for general plain text, but som…