2 papers
eess.AS2024
Lightweight Audio Segmentation for Long-form Speech Translation
Jaesong Lee, Soyoon Kim, Hanbyul Kim +1
Speech segmentation is an essential part of speech translation (ST) systems in real-world scenarios. Since most ST models are designed to process speech segments, long-form audio m…
cs.CL2023
I3D: Transformer architectures with input-dependent dynamic depth for speech recognition
Yifan Peng, Jaesong Lee, Shinji Watanabe
Transformer-based end-to-end speech recognition has achieved great success. However, the large footprint and computational overhead make it difficult to deploy these models in some…