2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2025★ 2 cited
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
Guangzhi Sun, Yudong Yang, Jimin Zhuang +5
While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been li…
eess.AS2024
Automatic Assessment of Dysarthria Using Audio-visual Vowel Graph Attention Network
Xiaokang Liu, Xiaoxia Du, Juan Liu +7
Automatic assessment of dysarthria remains a highly challenging task due to high variability in acoustic signals and the limited data. Currently, research on the automatic assessme…