1 citations · 1 across the 1 of their papers we have counts for
Showing 2024Show all
3 papers · 1 filter
cs.CV2024
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Guangzhi Sun, Wenyi Yu, Changli Tang +7
Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper propo…
physics.optics2024
Theoretical efficiency limit of diffractive input couplers in augmented reality waveguides
Zhexin Zhao, Yun-Han Lee, Xiayu Feng +3
Considerable efforts have been devoted into augmented reality (AR) displays to enable the immersive user experience in the wearable glasses form factor. Transparent waveguide combi…
eess.AS2024★ 1 cited
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model
Dongdi Zhao, Jianbo Ma, Lu Lu +6
Far-field speech recognition is a challenging task that conventionally uses signal processing beamforming to attack noise and interference problem. But the performance has been fou…