9 citations · 40 across the 10 of their papers we have counts for
5 papers · 1 filter
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Zihang Dai, Zhilin Yang, Yiming Yang +3
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architect…
Characterizing and Avoiding Negative Transfer
Zirui Wang, Zihang Dai, Barnabás Póczos +1
When labeled data is scarce for a specific target task, transfer learning often offers an effective solution by utilizing data from a related source task. However, when transferrin…
Towards more Reliable Transfer Learning
Zirui Wang, Jaime Carbonell
Multi-source transfer learning has been proven effective when within-target labeled data is scarce. Previous work focuses primarily on exploiting domain similarities and assumes th…
Nonparametric Neural Networks
George Philipp, Jaime G. Carbonell
Automatically determining the optimal size of a neural network for a given task without prior information currently requires an expensive global search and training many networks f…
Bounds on the Minimax Rate for Estimating a Prior over a VC Class from Independent Learning Tasks
Liu Yang, Steve Hanneke, Jaime Carbonell
We study the optimal rates of convergence for estimating a prior distribution over a VC class from a sequence of independent data sets respectively labeled by independent target fu…