1 paper · 1 filter
Rui Hu, Delai Qiu, Yining Wang +2
Omni-modal large language models (OLLMs) offer a promising end-to-end solution for slide-enhanced speech recognition due to their inherent multimodal capabilities. However, we foun…