4 citations · 8 across the 6 of their papers we have counts for
5 papers
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
Kwanyong Park, Kuniaki Saito, Donghyun Kim
Vision-language (VL) models often exhibit a limited understanding of complex expressions of visual objects (e.g., attributes, shapes, and their relations), given complex and divers…
Mind the Backbone: Minimizing Backbone Distortion for Robust Object Detection
Kuniaki Saito, Donghyun Kim, Piotr Teterwak +2
Building object detectors that are robust to domain shifts is critical for real-world applications. Prior approaches fine-tune a pre-trained backbone and risk overfitting it to in-…
Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval
Kuniaki Saito, Kihyuk Sohn, Xiang Zhang +4
In Composed Image Retrieval (CIR), a user combines a query image with text to describe their intended target. Existing methods rely on supervised learning of CIR models using label…
Learning to Detect Every Thing in an Open World
Kuniaki Saito, Ping Hu, Trevor Darrell +1
Many open-world applications require the detection of novel objects, yet state-of-the-art object detection and instance segmentation networks do not excel at this task. The key iss…
DeMIAN: Deep Modality Invariant Adversarial Network
Kuniaki Saito, Yusuke Mukuta, Yoshitaka Ushiku +1
Obtaining common representations from different modalities is important in that they are interchangeable with each other in a classification problem. For example, we can train a cl…