33 citations · 34 across the 4 of their papers we have counts for
5 papers
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
Xianhang Li, Yanqing Liu, Haoqin Tu +2
OpenAI's CLIP, released in early 2021, have long been the go-to choice of vision encoder for building multimodal foundation models. Although recent alternatives such as SigLIP have…
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
Sucheng Ren, Hongru Zhu, Chen Wei +3
This paper presents a new self-supervised video representation learning framework, ARVideo, which autoregressively predicts the next video token in a tailored sequence order. Two k…
Revisiting Adversarial Training at Scale
Zeyu Wang, Xianhang Li, Hongru Zhu +1
The machine learning community has witnessed a drastic change in the training pipeline, pivoted by those ''foundation models'' with unprecedented scales. However, the field of adve…
Rejuvenating image-GPT as Strong Visual Representation Learners
Sucheng Ren, Zeyu Wang, Hongru Zhu +3
This paper enhances image-GPT (iGPT), one of the pioneering works that introduce autoregressive pretraining to predict the next pixels for visual representation learning. Two simpl…
Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models
Hongru Zhu, Peng Tang, Jeongho Park +2
Most objects in the visual world are partially occluded, but humans can recognize them without difficulty. However, it remains unknown whether object recognition models like convol…