53 citations · 168 across the 17 of their papers we have counts for
6 papers · 1 filter
KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models
Zhiwei Jia, Pradyumna Narayana, Arjun R. Akula +4
Image ad understanding is a crucial task with wide real-world applications. Although highly challenging with the involvement of diverse atypical scenes, real-world entities, and re…
ActiveZero: Mixed Domain Learning for Active Stereovision with Zero Annotation
Isabella Liu, Edward Yang, Jianyu Tao +5
Traditional depth sensors generate accurate real world depth estimates that surpass even the most advanced learning approaches trained only on simulation domains. Since ground trut…
FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization
Xingchao Liu, Chengyue Gong, Lemeng Wu +3
Generating images from natural language instructions is an intriguing yet highly challenging task. We approach text-to-image generation by combining the power of the retrained CLIP…
Beyond Holistic Object Recognition: Enriching Image Understanding with Part States
Cewu Lu, Hao Su, Yongyi Lu +3
Important high-level vision tasks such as human-object interaction, image captioning and robotic manipulation require rich semantic descriptions of objects at part level. Based upo…
3D-Assisted Image Feature Synthesis for Novel Views of an Object
Hao Su, Fan Wang, Li Yi +1
Comparing two images in a view-invariant way has been a challenging problem in computer vision for a long time, as visual features are not stable under large view point changes. In…
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su +9
The ImageNet Large Scale Visual Recognition Challenge is a benchmark in object category classification and detection on hundreds of object categories and millions of images. The ch…