13 citations · 13 across the 3 of their papers we have counts for
3 papers
cs.CL2023
Principled Gradient-based Markov Chain Monte Carlo for Text Generation
Li Du, Afra Amini, Lucas Torroba Hennigen +4
Recent papers have demonstrated the possibility of energy-based text generation by adapting gradient-based sampling algorithms, a paradigm of MCMC algorithms that promises fast con…
cs.CL2023
Deriving Language Models from Masked Language Models
Lucas Torroba Hennigen, Yoon Kim
Masked language models (MLM) do not explicitly define a distribution over language, i.e., they are not language models per se. However, recent work has implicitly treated them as s…
cs.LG2023★ 13 cited
Learning to Grow Pretrained Models for Efficient Transformer Training
Peihao Wang, Rameswar Panda, Lucas Torroba Hennigen +6
Scaling transformers has led to significant breakthroughs in many domains, leading to a paradigm in which larger versions of existing models are trained and released on a periodic…