7 papers · 1 filter
A-SelecT: Automatic Timestep Selection for Diffusion Transformer Representation Learning
Changyu Liu, James Chenhao Liang, Wenhao Yang +6
Diffusion models have significantly reshaped the field of generative artificial intelligence and are now increasingly explored for their capacity in discriminative representation l…
Visual Fourier Prompt Tuning
Runjia Zeng, Cheng Han, Qifan Wang +5
With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visu…
AMD: Automatic Multi-step Distillation of Large-scale Vision Models
Cheng Han, Qifan Wang, Sohail A. Dianat +6
Transformer-based architectures have become the de-facto standard models for diverse vision tasks owing to their superior performance. As the size of the models continues to scale…
Self-supervised Adversarial Training of Monocular Depth Estimation against Physical-World Attacks
Zhiyuan Cheng, Cheng Han, James Liang +3
Monocular Depth Estimation (MDE) plays a vital role in applications such as autonomous driving. However, various attacks target MDE models, with physical attacks posing significant…
ProMotion: Prototypes As Motion Learners
Yawen Lu, Dongfang Liu, Qifan Wang +6
In this work, we introduce ProMotion, a unified prototypical framework engineered to model fundamental motion tasks. ProMotion offers a range of compelling attributes that set it a…
Prototypical Transformer as Unified Motion Learners
Cheng Han, Yawen Lu, Guohao Sun +9
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoForme…