13 citations · 23 across the 12 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 5 cited
VITA: Towards Open-Source Interactive Omni Multimodal LLM
Chaoyou Fu, Haojia Lin, Zuwei Long +16
The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas.…
cs.CV2023★ 13 cited
A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Chaoyou Fu, Renrui Zhang, Zihan Wang +15
The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow L…