5 citations · 11 across the 5 of their papers we have counts for
5 papers
Semantic Alignment for Multimodal Large Language Models
Tao Wu, Mengze Li, Jingyuan Chen +6
Research on Multi-modal Large Language Models (MLLMs) towards the multi-image cross-modal instruction has received increasing attention and made significant progress, particularly…
Enabling Collaborative Clinical Diagnosis of Infectious Keratitis by Integrating Expert Knowledge and Interpretable Data-driven Intelligence
Zhengqing Fang, Shuowen Zhou, Zhouhang Yuan +9
Although data-driven artificial intelligence (AI) in medical image diagnosis has shown impressive performance in silico, the lack of interpretability makes it difficult to incorpor…
HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object Grounding
Mengze Li, Tianbao Wang, Haoyu Zhang +6
Video Object Grounding (VOG) is the problem of associating spatial object regions in the video to a descriptive natural language query. This is a challenging vision-language task t…
BOSS: Bottom-up Cross-modal Semantic Composition with Hybrid Counterfactual Training for Robust Content-based Image Retrieval
Wenqiao Zhang, Jiannan Guo, Mengze Li +5
Content-Based Image Retrieval (CIR) aims to search for a target image by concurrently comprehending the composition of an example image and a complementary text, which potentially…
Enhancing Fairness of Visual Attribute Predictors
Tobias Hänel, Nishant Kumar, Dmitrij Schlesinger +4
The performance of deep neural networks for image recognition tasks such as predicting a smiling face is known to degrade with under-represented classes of sensitive attributes. We…