4 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
Hang Hua, Yunlong Tang, Ziyun Zeng +5
The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal understanding, enabling more sophisticated and accurate integration of visual and textual in…
cs.CV2023★ 4 cited
GPT-4V(ision) as A Social Media Analysis Engine
Hanjia Lyu, Jinfa Huang, Daoan Zhang +6
Recent research has offered insights into the extraordinary capabilities of Large Multimodal Models (LMMs) in various general vision and language tasks. There is growing interest i…