activity
20192026
most citedLearning Low Dimensional State Spaces with Overparameterized Recurrent Neural Nets

2 citations · 4 across the 6 of their papers we have counts for

collaborators

11 papers

cs.CL2026

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik +2

Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: n…

cs.LG2024

Provable Benefits of Complex Parameterizations for Structured State Space Models

Yuval Ran-Milo, Eden Lumbroso, Edo Cohen-Karlik +3

Structured state space models (SSMs), the core engine behind prominent neural networks such as S4 and Mamba, are linear dynamical systems adhering to a specified structure, most no…

cs.SI2024★ 1 cited

Overcoming Order in Autoregressive Graph Generation

Edo Cohen-Karlik, Eyal Rozenberg, Daniel Freedman

Graph generation is a fundamental problem in various domains, including chemistry and social networks. Recent work has shown that molecular graph generation using recurrent neural…

cs.LG2024

Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial States

Noam Razin, Yotam Alexander, Edo Cohen-Karlik +3

In modern machine learning, models can often fit training data in numerous ways, some of which perform well on unseen (test) data, while others do not. Remarkably, in such cases gr…

cs.LG2022★ 2 cited

Learning Low Dimensional State Spaces with Overparameterized Recurrent Neural Nets

Edo Cohen-Karlik, Itamar Menuhin-Gruman, Raja Giryes +2

Overparameterization in deep learning typically refers to settings where a trained neural network (NN) has representational capacity to fit the training data in many ways, some of…

cs.LG2022

On the Implicit Bias of Gradient Descent for Temporal Extrapolation

Edo Cohen-Karlik, Avichai Ben David, Nadav Cohen +1

When using recurrent neural networks (RNNs) it is common practice to apply trained models to sequences longer than those seen in training. This "extrapolating" usage deviates from…