2 citations · 3 across the 3 of their papers we have counts for
4 papers
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
Michael R. Metel, Peng Lu, Boxing Chen +2
We present a simple on the fly method for faster inference of large language models. Unlike other (self-)speculative decoding techniques, our method does not require fine-tuning or…
Hyperparameter Optimization for Large Language Model Instruction-Tuning
Christophe Tribes, Sacha Benarroch-Lelong, Peng Lu +1
The fine-tuning of Large Language Models (LLMs) has enabled them to recently achieve milestones in natural language processing applications. The emergence of ever larger LLMs has p…
Mathematical Challenges in Deep Learning
Vahid Partovi Nia, Guojun Zhang, Ivan Kobyzev +8
Deep models are dominating the artificial intelligence (AI) industry since the ImageNet challenge in 2012. The size of deep models is increasing ever since, which brings new challe…
Learning Functions on Multiple Sets using Multi-Set Transformers
Kira Selby, Ahmad Rashid, Ivan Kobyzev +2
We propose a general deep architecture for learning functions on multiple permutation-invariant sets. We also show how to generalize this architecture to sets of elements of any di…