activity
20202025
most citedTemporal Context Aggregation for Video Retrieval with Contrastive Learning

8 citations · 35 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2025

A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning

Xin Wen, Bingchen Zhao, Yilun Chen +2

Pre-trained vision models (PVMs) are fundamental to modern robotics, yet their optimal configuration remains unclear. Through systematic evaluation, we find that while DINO and iBO…

cs.CV2025

"Principal Components" Enable A New Language of Images

Xin Wen, Bingchen Zhao, Ismail Elezi +2

We introduce a novel visual tokenization framework that embeds a provable PCA-like structure into the latent token space. While existing visual tokenizers primarily optimize for re…

cs.CV2024

What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights

Xin Wen, Bingchen Zhao, Yilun Chen +2

Severe data imbalance naturally exists among web-scale vision-language datasets. Despite this, we find CLIP pre-trained thereupon exhibits notable robustness to the data imbalance…

cs.CV2023★ 2 cited

Learning Semi-supervised Gaussian Mixture Models for Generalized Category Discovery

Bingchen Zhao, Xin Wen, Kai Han

In this paper, we address the problem of generalized category discovery (GCD), \ie, given a set of images where part of them are labelled and the rest are not, the task is to autom…

cs.CV2023★ 2 cited

Incremental Generalized Category Discovery

Bingchen Zhao, Oisin Mac Aodha

We explore the problem of Incremental Generalized Category Discovery (IGCD). This is a challenging category incremental learning setting where the goal is to develop models that ca…

cs.CV2023

OOD-CV-v2: An extended Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images

Bingchen Zhao, Jiahao Wang, Wufei Ma +7

Enhancing the robustness of vision algorithms in real-world scenarios is challenging. One reason is that existing robustness benchmarks are limited, as they either rely on syntheti…