Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Training and Inference Efficiency of Encoder-Decoder Speech Models
Piotr Żelasko, Kunal Dhawan, Daniel Galvez +7
Attention encoder-decoder model architecture is the backbone of several recent top performing foundation speech models: Whisper, Seamless, OWSM, and Canary-1B. However, the reporte…
cs.CL2024
EMMeTT: Efficient Multimodal Machine Translation Training
Piotr Żelasko, Zhehuai Chen, Mengru Wang +7
A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses…