Publications (12)
Correlation Alignment for Unsupervised Domain Adaptation
Baochen Sun, Jiashi Feng, Kate Saenko
In this chapter, we present CORrelation ALignment (CORAL), a simple yet effective method for unsupervised domain adaptation. CORAL minimizes domain shift by aligning the second-ord…
LOWA: Localize Objects in the Wild with Attributes
Xiaoyuan Guo, Kezhen Chen, Jinmeng Rao +3
We present LOWA, a novel method for localizing objects with attributes effectively in the wild. It aims to address the insufficiency of current open-vocabulary object detectors, wh…
Deep CORAL: Correlation Alignment for Deep Domain Adaptation
Baochen Sun, Kate Saenko
Deep neural networks are able to learn powerful representations from large quantities of labeled input data, however they cannot always generalize well across changes in input dist…
Learning Deep Object Detectors from 3D Models
Xingchao Peng, Baochen Sun, Karim Ali +1
Crowdsourced 3D CAD models are becoming easily accessible online, and can potentially generate an infinite number of training images for almost any object category.We show that aug…
What Do Deep CNNs Learn About Objects?
Xingchao Peng, Baochen Sun, Karim Ali +1
Deep convolutional neural networks learn extremely powerful image representations, yet most of that power is hidden in the millions of deep-layer parameters. What exactly do these…
A Fast Minimization Algorithm for the Euler Elastica Model Based on a Bilinear Decomposition
Zhifang Liu, Baochen Sun, Xue-Cheng Tai +2
The Euler Elastica (EE) model with surface curvature can generate artifact-free results compared with the traditional total variation regularization model in image processing. Howe…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and Beyond
Thanh Vu, Baochen Sun, Bodi Yuan +3
The success of data mixing augmentations in image classification tasks has been well-received. However, these techniques cannot be readily applied to object detection due to challe…
Return of Frustratingly Easy Domain Adaptation
Baochen Sun, Jiashi Feng, Kate Saenko
Unlike human learning, machine learning often fails to handle changes between training (source) and test (target) input distributions. Such domain shifts, common in practical scena…
Modeling Radiometric Uncertainty for Vision with Tone-mapped Color Images
Ayan Chakrabarti, Ying Xiong, Baochen Sun +4
To produce images that are suitable for display, tone-mapping is widely used in digital cameras to map linear color measurements into narrow gamuts with limited dynamic range. This…
Higher Layers Need More LoRA Experts
Chongyang Gao, Kezhen Chen, Jinmeng Rao +7
Parameter-efficient tuning (PEFT) techniques like low-rank adaptation (LoRA) offer training efficiency on Large Language Models, but their impact on model performance remains limit…
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
Jiaying Lu, Jinmeng Rao, Kezhen Chen +5
Large Vision-Language Models (LVLMs) offer remarkable benefits for a variety of vision-language tasks. However, a challenge hindering their application in real-world scenarios, par…