153 citations · 197 across the 9 of their papers we have counts for
17 papers
Multitask Vision-Language Prompt Tuning
Sheng Shen, Shijia Yang, Tianjun Zhang +4
Prompt Tuning, conditioning on task-specific learned prompt vectors, has emerged as a data-efficient and parameter-efficient method for adapting large pretrained vision-language mo…
What Language Model to Train if You Have One Million GPU Hours?
Teven Le Scao, Thomas Wang, Daniel Hesslow +16
The crystallization of modeling methods around the Transformer architecture has been a boon for practitioners. Simple, well-motivated architectural variations can transfer across t…
ITSRN++: Stronger and Better Implicit Transformer Network for Continuous Screen Content Image Super-Resolution
Sheng Shen, Huanjing Yue, Jingyu Yang +1
Nowadays, online screen sharing and remote cooperation are becoming ubiquitous. However, the screen content may be downsampled and compressed during transmission, while it may be d…
One Parameter Defense -- Defending against Data Inference Attacks via Differential Privacy
Dayong Ye, Sheng Shen, Tianqing Zhu +2
Machine learning models are vulnerable to data inference attacks, such as membership inference and model inversion attacks. In these types of breaches, an adversary attempts to inf…
Staged Training for Transformer Language Models
Sheng Shen, Pete Walsh, Kurt Keutzer +3
The current standard approach to scaling transformer language models trains each model size from a different random initialization. As an alternative, we consider a staged training…
What's Hidden in a One-layer Randomly Weighted Transformer?
Sheng Shen, Zhewei Yao, Douwe Kiela +2
We demonstrate that, hidden within one-layer randomly weighted neural networks, there exist subnetworks that can achieve impressive performance, without ever modifying the weight i…