papers

Publications (8)

cs.LG2022

Implicit Regularization Towards Rank Minimization in ReLU Networks

Nadav Timor, Gal Vardi, Ohad Shamir

We study the conjectured relationship between the implicit regularization in neural networks, trained with gradient-based methods, and rank minimization of their weight matrices. P…

cs.CL2023

StarCoder: may the source be with you!

Raymond Li, Loubna Ben Allal, Yangtian Zi +64

The BigCode community, an open-scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder and StarCoderBase…

cs.CL2025

Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies

Nadav Timor, Jonathan Mamou, Daniel Korat +5

Accelerating the inference of large language models (LLMs) is a critical challenge in generative AI. Speculative decoding (SD) methods offer substantial efficiency gains by generat…

cs.CL2024

Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models

Jonathan Mamou, Oren Pereg, Daniel Korat +4

Speculative decoding is commonly used for reducing the inference latency of large language models. Its effectiveness depends highly on the speculation lookahead (SL)-the number of…

cs.LG2025

NdLinear: Preserving Multi-Dimensional Structure for Parameter-Efficient Neural Networks

Alex Reneau, Jerry Yao-Chieh Hu, Zhongfang Zhuang +6

In deep learning, processing multidimensional inputs (e.g., images, medical scans, and time series) is an important task that often requires flattening the inputs. We introduce $\m…

cs.LG2025

Out-of-Vocabulary Sampling Boosts Speculative Decoding

Nadav Timor, Jonathan Mamou, Oren Pereg +2

Speculative decoding relies on fast and accurate drafters. Recent state-of-the-art language models employ larger and larger vocabularies, which significantly slows down drafters. O…

cs.LG2026

On Training in Imagination

Nadav Timor, Ravid Shwartz-Ziv, Micah Goldblum +2

State-of-the-art model-based reinforcement learning methods train policies on imagined rollouts. These rollouts are trajectories generated by a learned dynamics model and are score…

cs.DC2025

Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference

Nadav Timor, Jonathan Mamou, Daniel Korat +6

This paper introduces distributed speculative inference (DSI), a novel inference algorithm that is provably faster than speculative inference (SI) [leviathan2023, chen2023, miao202…