28 citations · 35 across the 19 of their papers we have counts for
5 papers · 1 filter
Creating a Lens of Chinese Culture: A Multimodal Dataset for Chinese Pun Rebus Art Understanding
Tuo Zhang, Tiantian Feng, Yibin Ni +7
Large vision-language models (VLMs) have demonstrated remarkable abilities in understanding everyday content. However, their performance in the domain of art, particularly cultural…
ConPro: Learning Severity Representation for Medical Images using Contrastive Learning and Preference Optimization
Hong Nguyen, Hoang Nguyen, Melinda Chang +3
Understanding the severity of conditions shown in images in medical diagnosis is crucial, serving as a key guide for clinical assessment, treatment, as well as evaluating longitudi…
Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?
Tiantian Feng, Daniel Yang, Digbalay Bose +1
Multi-modal learning has emerged as an increasingly promising avenue in vision recognition, driving innovations across diverse domains ranging from media and education to healthcar…
MM-AU:Towards Multimodal Understanding of Advertisement Videos
Digbalay Bose, Rajat Hebbar, Tiantian Feng +3
Advertisement videos (ads) play an integral part in the domain of Internet e-commerce as they amplify the reach of particular products to a broad audience or can serve as a medium…
Contextually-rich human affect perception using multimodal scene information
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli +1
The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect percept…