26 citations · 44 across the 5 of their papers we have counts for
9 papers
Is the Number of Trainable Parameters All That Actually Matters?
Amélie Chatelain, Amine Djeghri, Daniel Hesslow +2
Recent work has identified simple empirical scaling laws for language models, linking compute budget, dataset size, model size, and autoregressive modeling loss. The validity of th…
ROPUST: Improving Robustness through Fine-tuning with Photonic Processors and Synthetic Gradients
Alessandro Cappelli, Julien Launay, Laurent Meunier +2
Robustness to adversarial attacks is typically obtained through expensive adversarial training with Projected Gradient Descent. Here we introduce ROPUST, a remarkably simple and ef…
Photonic co-processors in HPC: using LightOn OPUs for Randomized Numerical Linear Algebra
Daniel Hesslow, Alessandro Cappelli, Igor Carron +6
Randomized Numerical Linear Algebra (RandNLA) is a powerful class of methods, widely used in High Performance Computing (HPC). RandNLA provides approximate solutions to linear alge…
Contrastive Embeddings for Neural Architectures
Daniel Hesslow, Iacopo Poli
The performance of algorithms for neural architecture search strongly depends on the parametrization of the search space. We use contrastive learning to identify networks across di…
Hardware Beyond Backpropagation: a Photonic Co-Processor for Direct Feedback Alignment
Julien Launay, Iacopo Poli, Kilian Müller +5
The scaling hypothesis motivates the expansion of models past trillions of parameters as a path towards better performance. Recent significant developments, such as GPT-3, have bee…
Light-in-the-loop: using a photonics co-processor for scalable training of neural networks
Julien Launay, Iacopo Poli, Kilian Müller +4
As neural networks grow larger and more complex and data-hungry, training costs are skyrocketing. Especially when lifelong learning is necessary, such as in recommender systems or…