1 paper
Yichen Lu, Jiaqi Song, Xuankai Chang +3
In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aid…