papers

Publications (12)

cs.CV2016

Correlation Alignment for Unsupervised Domain Adaptation

Baochen Sun, Jiashi Feng, Kate Saenko

In this chapter, we present CORrelation ALignment (CORAL), a simple yet effective method for unsupervised domain adaptation. CORAL minimizes domain shift by aligning the second-ord…

cs.CV2023

LOWA: Localize Objects in the Wild with Attributes

Xiaoyuan Guo, Kezhen Chen, Jinmeng Rao +3

We present LOWA, a novel method for localizing objects with attributes effectively in the wild. It aims to address the insufficiency of current open-vocabulary object detectors, wh…

cs.CV2016

Deep CORAL: Correlation Alignment for Deep Domain Adaptation

Baochen Sun, Kate Saenko

Deep neural networks are able to learn powerful representations from large quantities of labeled input data, however they cannot always generalize well across changes in input dist…

cs.CV2015

Learning Deep Object Detectors from 3D Models

Xingchao Peng, Baochen Sun, Karim Ali +1

Crowdsourced 3D CAD models are becoming easily accessible online, and can potentially generate an infinite number of training images for almost any object category.We show that aug…

cs.CV2015

What Do Deep CNNs Learn About Objects?

Xingchao Peng, Baochen Sun, Karim Ali +1

Deep convolutional neural networks learn extremely powerful image representations, yet most of that power is hidden in the millions of deep-layer parameters. What exactly do these…

math.OC2023

A Fast Minimization Algorithm for the Euler Elastica Model Based on a Bilinear Decomposition

Zhifang Liu, Baochen Sun, Xue-Cheng Tai +2

The Euler Elastica (EE) model with surface curvature can generate artifact-free results compared with the traditional total variation regularization model in image processing. Howe…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

cs.CV2023

Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and Beyond

Thanh Vu, Baochen Sun, Bodi Yuan +3

The success of data mixing augmentations in image classification tasks has been well-received. However, these techniques cannot be readily applied to object detection due to challe…

cs.CV2015

Return of Frustratingly Easy Domain Adaptation

Baochen Sun, Jiashi Feng, Kate Saenko

Unlike human learning, machine learning often fails to handle changes between training (source) and test (target) input distributions. Such domain shifts, common in practical scena…

cs.CV2014

Modeling Radiometric Uncertainty for Vision with Tone-mapped Color Images

Ayan Chakrabarti, Ying Xiong, Baochen Sun +4

To produce images that are suitable for display, tone-mapping is widely used in digital cameras to map linear color measurements into narrow gamuts with limited dynamic range. This…

cs.CL2024

Higher Layers Need More LoRA Experts

Chongyang Gao, Kezhen Chen, Jinmeng Rao +7

Parameter-efficient tuning (PEFT) techniques like low-rank adaptation (LoRA) offer training efficiency on Large Language Models, but their impact on model performance remains limit…

cs.CV2024

Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models

Jiaying Lu, Jinmeng Rao, Kezhen Chen +5

Large Vision-Language Models (LVLMs) offer remarkable benefits for a variety of vision-language tasks. However, a challenge hindering their application in real-world scenarios, par…