activity
20202026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Bingyi Cao, Koert Chen, Kevis-Kokitsi Maninis +16

Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation…

cs.CV2024

Learning Visual Composition through Improved Semantic Guidance

Austin Stone, Hagen Soltau, Robert Geirhos +6

Visual imagery does not consist of solitary objects, but instead reflects the composition of a multitude of fluid concepts. While there have been great advances in visual represent…

cs.CV2024

TIPS: Text-Image Pretraining with Spatial awareness

Kevis-Kokitsi Maninis, Kaifeng Chen, Soham Ghosh +11

While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense und…

cs.CV2022

Semantic-shape Adaptive Feature Modulation for Semantic Image Synthesis

Zhengyao Lv, Xiaoming Li, Zhenxing Niu +2

Recent years have witnessed substantial progress in semantic image synthesis, it is still challenging in synthesizing photo-realistic images with rich details. Most previous method…

cs.CV2020

Google Landmarks Dataset v2 -- A Large-Scale Benchmark for Instance-Level Recognition and Retrieval

Tobias Weyand, Andre Araujo, Bingyi Cao +1

While image retrieval and instance recognition techniques are progressing rapidly, there is a need for challenging datasets to accurately measure their performance -- while posing…

cs.CV2020

Unifying Deep Local and Global Features for Image Search

Bingyi Cao, Andre Araujo, Jack Sim

Image retrieval is the problem of searching an image database for items that are similar to a query image. To address this task, two main types of image representations have been s…