23 citations · 24 across the 2 of their papers we have counts for
4 papers
What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?
Thomas Wang, Adam Roberts, Daniel Hesslow +5
Large pretrained Transformer language models have been shown to exhibit zero-shot generalization, i.e. they can perform a wide variety of tasks that they were not explicitly traine…
Is the Number of Trainable Parameters All That Actually Matters?
Amélie Chatelain, Amine Djeghri, Daniel Hesslow +2
Recent work has identified simple empirical scaling laws for language models, linking compute budget, dataset size, model size, and autoregressive modeling loss. The validity of th…
Photonic co-processors in HPC: using LightOn OPUs for Randomized Numerical Linear Algebra
Daniel Hesslow, Alessandro Cappelli, Igor Carron +6
Randomized Numerical Linear Algebra (RandNLA) is a powerful class of methods, widely used in High Performance Computing (HPC). RandNLA provides approximate solutions to linear alge…
Contrastive Embeddings for Neural Architectures
Daniel Hesslow, Iacopo Poli
The performance of algorithms for neural architecture search strongly depends on the parametrization of the search space. We use contrastive learning to identify networks across di…