4 papers · 1 filter
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
Shuai Zhao, Ruijie Quan, Linchao Zhu +1
Pre-trained vision-language models~(VLMs) are the de-facto foundation models for various downstream tasks. However, scene text recognition methods still prefer backbones pre-traine…
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
Xu Zhang, Zhedong Zheng, Linchao Zhu +1
Composed image retrieval extends content-based image retrieval systems by enabling users to search using reference images and captions that describe their intention. Despite great…
Slimmable Networks for Contrastive Self-supervised Learning
Shuai Zhao, Linchao Zhu, Xiaohan Wang +1
Self-supervised learning makes significant progress in pre-training large models, but struggles with small models. Mainstream solutions to this problem rely mainly on knowledge dis…
Scalable Video Object Segmentation with Identification Mechanism
Zongxin Yang, Jiaxu Miao, Yunchao Wei +3
This paper delves into the challenges of achieving scalable and effective multi-object modeling for semi-supervised Video Object Segmentation (VOS). Previous VOS methods decode fea…