1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
Ziyuan Huang, Kaixiang Ji, Biao Gong +6
This paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs). Our approach employs a sequence…
cs.CV2024★ 1 cited
Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification
Shulei Qiu, Wanqi Yang, Ming Yang
Our research focuses on few-shot fine-grained image classification, which faces two major challenges: appearance similarity of fine-grained objects and limited number of samples. T…
cs.CV2024
Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model
Kangyang Xie, Binbin Yang, Hao Chen +5
Beyond the superiority of the text-to-image diffusion model in generating high-quality images, recent studies have attempted to uncover its potential for adapting the learned seman…