11 citations · 49 across the 10 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 1 cited
Instruct-Imagen: Image Generation with Multi-modal Instruction
Hexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su +9
This paper presents instruct-imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce *multi-modal instruction* for image…
cs.CV2023★ 3 cited
Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities
Hexiang Hu, Yi Luan, Yang Chen +5
Large-scale multi-modal pre-training models such as CLIP and PaLI exhibit strong generalization on various visual domains and tasks. However, existing image classification benchmar…