activity
20172024
most citedOpen-Vocabulary Object Detection Using Captions

17 citations · 47 across the 9 of their papers we have counts for

collaborators

15 papers

cs.CV2024★ 1 cited

Imagine yourself: Tuning-Free Personalized Image Generation

Zecheng He, Bo Sun, Felix Juefei-Xu +14

Diffusion models have demonstrated remarkable efficacy across various image-to-image tasks. In this research, we introduce Imagine yourself, a state-of-the-art model designed for p…

cs.CV2023

Learning from Children: Improving Image-Caption Pretraining via Curriculum

Hammad A. Ayyubi, Rahul Lokesh, Alireza Zareian +2

Image-caption pretraining has been quite successfully used for downstream vision tasks like zero-shot image classification and object detection. However, image-caption pretraining…

cs.CV2022

GOCA: Guided Online Cluster Assignment for Self-Supervised Video Representation Learning

Huseyin Coskun, Alireza Zareian, Joshua L. Moore +2

Clustering is a ubiquitous tool in unsupervised learning. Most of the existing self-supervised representation learning methods typically cluster samples based on visually dominant…

cs.CV2021

SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense Reasoning

Zhecan Wang, Haoxuan You, Liunian Harold Li +5

Answering complex questions about images is an ambitious goal for machine intelligence, which requires a joint understanding of images, text, and commonsense knowledge, as well as…

cs.CV2020★ 17 cited

Open-Vocabulary Object Detection Using Captions

Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu +1

Despite the remarkable accuracy of deep neural networks in object detection, they are costly to train and scale due to supervision requirements. Particularly, learning more object…

cs.CL2020★ 6 cited

Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions

Liunian Harold Li, Haoxuan You, Zhecan Wang +3

Pre-trained contextual vision-and-language (V&L) models have achieved impressive performance on various benchmarks. However, existing models require a large amount of parallel imag…