activity
20162025
most citedDeMIAN: Deep Modality Invariant Adversarial Network

4 citations · 8 across the 6 of their papers we have counts for

collaborators

5 papers

cs.CV2024

Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection

Kwanyong Park, Kuniaki Saito, Donghyun Kim

Vision-language (VL) models often exhibit a limited understanding of complex expressions of visual objects (e.g., attributes, shapes, and their relations), given complex and divers…

cs.CV20231 cited

Mind the Backbone: Minimizing Backbone Distortion for Robust Object Detection

Kuniaki Saito, Donghyun Kim, Piotr Teterwak +2

Building object detectors that are robust to domain shifts is critical for real-world applications. Prior approaches fine-tune a pre-trained backbone and risk overfitting it to in-…

cs.CV20232 cited

Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval

Kuniaki Saito, Kihyuk Sohn, Xiang Zhang +4

In Composed Image Retrieval (CIR), a user combines a query image with text to describe their intended target. Existing methods rely on supervised learning of CIR models using label…

cs.CV20211 cited

Learning to Detect Every Thing in an Open World

Kuniaki Saito, Ping Hu, Trevor Darrell +1

Many open-world applications require the detection of novel objects, yet state-of-the-art object detection and instance segmentation networks do not excel at this task. The key iss…

cs.LG20164 cited

DeMIAN: Deep Modality Invariant Adversarial Network

Kuniaki Saito, Yusuke Mukuta, Yoshitaka Ushiku +1

Obtaining common representations from different modalities is important in that they are interchangeable with each other in a classification problem. For example, we can train a cl…