26 citations · 68 across the 9 of their papers we have counts for
7 papers · 1 filter
Modulating Pretrained Diffusion Models for Multimodal Image Synthesis
Cusuh Ham, James Hays, Jingwan Lu +3
We present multimodal conditioning modules (MCM) for enabling conditional image synthesis using pretrained diffusion models. Previous multimodal synthesis works rely on training ne…
A Sketch Is Worth a Thousand Words: Image Retrieval with Text and Sketch
Patsorn Sangkloy, Wittawat Jitkrittum, Diyi Yang +1
We address the problem of retrieving images with both a sketch and a text query. We present TASK-former (Text And SKetch transformer), an end-to-end trainable model for image retri…
MSeg: A Composite Dataset for Multi-domain Semantic Segmentation
John Lambert, Zhuang Liu, Ozan Sener +2
We present MSeg, a composite dataset that unifies semantic segmentation datasets from different domains. A naive merge of the constituent datasets yields poor performance due to in…
Complex Event Recognition from Images with Few Training Examples
Unaiza Ahsan, Chen Sun, James Hays +1
We propose to leverage concept-level representations for complex event recognition in photographs given limited training examples. We introduce a novel framework to discover event…
Scribbler: Controlling Deep Image Synthesis with Sketch and Color
Patsorn Sangkloy, Jingwan Lu, Chen Fang +2
Recently, there have been several promising methods to generate realistic imagery from deep convolutional networks. These methods sidestep the traditional computer graphics renderi…
StuffNet: Using 'Stuff' to Improve Object Detection
Samarth Brahmbhatt, Henrik I. Christensen, James Hays
We propose a Convolutional Neural Network (CNN) based algorithm - StuffNet - for object detection. In addition to the standard convolutional features trained for region proposal an…