activity
20162026
most citedVisualizing Dynamics: from t-SNE to SEMI-MDPs

12 citations · 51 across the 21 of their papers we have counts for

collaborators
Showing 2019Show all

7 papers · 1 filter

physics.optics2019

Deep learning reconstruction of ultrashort pulses from 2D spatial intensity patterns recorded by an all-in-line system in a single-shot

Ron Ziv, Alex Dikopoltsev, Tom Zahavy +4

We propose a simple all-in-line single-shot scheme for diagnostics of ultrashort laser pulses, consisting of a multi-mode fiber, a nonlinear crystal and a CCD camera. The system re…

cs.LG2019

Apprenticeship Learning via Frank-Wolfe

Tom Zahavy, Alon Cohen, Haim Kaplan +1

We consider the applications of the Frank-Wolfe (FW) algorithm for Apprenticeship Learning (AL). In this setting, we are given a Markov Decision Process (MDP) without an explicit r…

cs.LG2019

Inverse Reinforcement Learning in Contextual MDPs

Stav Belogolovsky, Philip Korsunsky, Shie Mannor +2

We consider the task of Inverse Reinforcement Learning in Contextual Markov Decision Processes (MDPs). In this setting, contexts, which define the reward and transition kernel, are…

cs.LG2019

Unknown mixing times in apprenticeship and reinforcement learning

Tom Zahavy, Alon Cohen, Haim Kaplan +1

We derive and analyze learning algorithms for apprenticeship learning, policy evaluation, and policy gradient for average reward criteria. Existing algorithms explicitly require an…

cs.LG2019

Action Assembly: Sparse Imitation Learning for Text Based Games with Combinatorial Action Spaces

Chen Tessler, Tom Zahavy, Deborah Cohen +2

We propose a computationally efficient algorithm that combines compressed sensing with imitation learning to solve text-based games with combinatorial action spaces. Specifically,…

cs.LG2019

Planning in Hierarchical Reinforcement Learning: Guarantees for Using Local Policies

Tom Zahavy, Avinatan Hasidim, Haim Kaplan +1

We consider a settings of hierarchical reinforcement learning, in which the reward is a sum of components. For each component we are given a policy that maximizes it and our goal i…