Publications (19)
Visualizing How Embeddings Generalize
Xiaotong Liu, Hong Xuan, Zeyu Zhang +2
Deep metric learning is often used to learn an embedding function that captures the semantic differences within a dataset. A key factor in many problem domains is how this embeddin…
Deep Randomized Ensembles for Metric Learning
Hong Xuan, Richard Souvenir, Robert Pless
Learning embedding functions, which map semantically related inputs to nearby locations in a feature space supports a variety of classification and information retrieval tasks. In…
Hard negative examples are hard, but useful
Hong Xuan, Abby Stylianou, Xiaotong Liu +1
Triplet loss is an extremely common approach to distance metric learning. Representations of images from the same class are optimized to be mapped closer together in an embedding s…
Learning Geo-Temporal Image Features
Menghua Zhai, Tawfiq Salem, Connor Greenwell +3
We propose to implicitly learn to extract geo-temporal image features, which are mid-level features related to when and where an image was captured, by explicitly optimizing for a…
ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval
Eric Xing, Pranavi Kolouju, Robert Pless +2
Composed image retrieval (CIR) is the task of retrieving a target image specified by a query image and a relative text that describes a semantic modification to the query image. Ex…
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
Pranavi Kolouju, Eric Xing, Robert Pless +2
Composed image retrieval (CIR) enables users to search images using a reference image combined with textual modifications. Recent advances in vision-language models have improved C…
Will It Zero-Shot?: Predicting Zero-Shot Classification Performance For Arbitrary Queries
Kevin Robbins, Xiaotong Liu, Yu Wu +4
Vision-Language Models like CLIP create aligned embedding spaces for text and images, making it possible for anyone to build a visual classifier by simply naming the classes they w…
Hotels-50K: A Global Hotel Recognition Dataset
Abby Stylianou, Hong Xuan, Maya Shende +3
Recognizing a hotel from an image of a hotel room is important for human trafficking investigations. Images directly link victims to places and can help verify where victims have b…
On Seeding Watermarks to Detect Verbatim LLM Copy-Paste Responses
Aizierjiang Aiersilan, Artin Yousefi, Robert Pless
Large language models (LLMs) have made fluent essay writing, code drafting, and quiz answering instantly available to students at every level, from secondary school through graduat…
Visualizing Deep Similarity Networks
Abby Stylianou, Richard Souvenir, Robert Pless
For convolutional neural network models that optimize an image embedding, we propose a method to highlight the regions of images that contribute most to pairwise similarity. This w…
Deep Feature Interpolation for Image Content Changes
Paul Upchurch, Jacob Gardner, Geoff Pleiss +4
We propose Deep Feature Interpolation (DFI), a new data-driven baseline for automatic high-resolution image transformation. As the name suggests, it relies only on simple linear in…
DCAP: Deep Cross Attentional Product Network for User Response Prediction
Zekai Chen, Fangtian Zhong, Zhumin Chen +3
User response prediction, which aims to predict the probability that a user will provide a predefined positive response in a given context such as clicking on an ad or purchasing a…
Shadow Estimation Method for "The Episolar Constraint: Monocular Shape from Shadow Correspondence"
Austin Abrams, Chris Hawley, Kylia Miskell +3
Recovering shadows is an important step for many vision algorithms. Current approaches that work with time-lapse sequences are limited to simple thresholding heuristics. We show th…
DP2-Pub: Differentially Private High-Dimensional Data Publication with Invariant Post Randomization
Honglu Jiang, Haotian Yu, Xiuzhen Cheng +3
A large amount of high-dimensional and heterogeneous data appear in practical applications, which are often published to third parties for data analysis, recommendations, targeted…
Dissecting the impact of different loss functions with gradient surgery
Hong Xuan, Robert Pless
Pair-wise loss is an approach to metric learning that learns a semantic embedding by optimizing a loss function that encourages images from the same semantic class to be mapped clo…
TraffickCam: Explainable Image Matching For Sex Trafficking Investigations
Abby Stylianou, Richard Souvenir, Robert Pless
Investigations of sex trafficking sometimes have access to photographs of victims in hotel rooms. These images directly link victims to places, which can help verify where victims…
QuARI: Query Adaptive Retrieval Improvement
Eric Xing, Abby Stylianou, Robert Pless +1
Massive-scale pretraining has made vision-language models increasingly popular for image-to-image and text-to-image retrieval across a broad collection of domains. However, these m…
Classification and Visualization of Genotype x Phenotype Interactions in Biomass Sorghum
Abby Stylianou, Robert Pless, Nadia Shakoor +1
We introduce a simple approach to understanding the relationship between single nucleotide polymorphisms (SNPs), or groups of related SNPs, and the phenotypes they control. The pip…
Improved Embeddings with Easy Positive Triplet Mining
Hong Xuan, Abby Stylianou, Robert Pless
Deep metric learning seeks to define an embedding where semantically similar images are embedded to nearby locations, and semantically dissimilar images are embedded to distant loc…