5 citations · 8 across the 3 of their papers we have counts for
3 papers
cs.MM2025
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Yong Ren, Chenxing Li, Le Xu +7
Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remain…
cs.AI2023★ 3 cited
Active Reasoning in an Open-World Environment
Manjie Xu, Guangyuan Jiang, Wei Liang +2
Recent advances in vision-language learning have achieved notable success on complete-information question-answering datasets through the integration of extensive world knowledge.…
cs.CL2023★ 5 cited
MEWL: Few-shot multimodal word learning with referential uncertainty
Guangyuan Jiang, Manjie Xu, Shiji Xin +4
Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This…