2 citations · 3 across the 2 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024★ 1 cited
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
Zhangwei Gao, Zhe Chen, Erfei Cui +12
Multimodal large language models (MLLMs) have demonstrated impressive performance in vision-language tasks across a broad spectrum of domains. However, the large model scale and as…
cs.CV2024
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Dongyang Liu, Renrui Zhang, Longtian Qiu +16
We propose SPHINX-X, an extensive Multimodality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX…
cs.CV2023
Towards the Unification of Generative and Discriminative Visual Foundation Model: A Survey
Xu Liu, Tong Zhou, Yuanxin Wang +7
The advent of foundation models, which are pre-trained on vast datasets, has ushered in a new era of computer vision, characterized by their robustness and remarkable zero-shot gen…