167 citations · 332 across the 29 of their papers we have counts for
11 papers · 1 filter
Point Cloud Recognition with Position-to-Structure Attention Transformers
Zheng Ding, James Hou, Zhuowen Tu
In this paper, we present Position-to-Structure Attention Transformers (PS-Former), a Transformer-based algorithm for 3D point cloud recognition. PS-Former deals with the challenge…
On the Feasibility of Cross-Task Transfer with Model-Based Reinforcement Learning
Yifan Xu, Nicklas Hansen, Zirui Wang +3
Reinforcement Learning (RL) algorithms can solve challenging control problems directly from image observations, but they often require millions of environment interactions to do so…
An In-depth Study of Stochastic Backpropagation
Jun Fang, Mingze Xu, Hao Chen +3
In this paper, we provide an in-depth study of Stochastic Backpropagation (SBP) when training deep neural networks for standard image classification and object detection tasks. Dur…
Semi-supervised Vision Transformers at Scale
Zhaowei Cai, Avinash Ravichandran, Paolo Favaro +5
We study semi-supervised learning (SSL) for vision transformers (ViT), an under-explored topic despite the wide adoption of the ViT architectures to different tasks. To tackle this…
Open-Vocabulary Universal Image Segmentation with MaskCLIP
Zheng Ding, Jieke Wang, Zhuowen Tu
In this paper, we tackle an emerging computer vision task, open-vocabulary universal image segmentation, that aims to perform semantic/instance/panoptic segmentation (background se…
The Geometry of Multilingual Language Model Representations
Tyler A. Chang, Zhuowen Tu, Benjamin K. Bergen
We assess how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language. Using XLM-R as…