collaborators

9 papers

cs.LG2026

Reevaluating Policy Gradient Methods for Imperfect-Information Games

Max Rudolph, Nathan Lichtle, Sobhan Mohammadpour +6

In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed nu…

cs.LG2026

Factored Latent Action World Models

Zizhao Wang, Chang Shi, Jiaheng Hu +4

Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions provide a natural interface for u…

cs.RO2026

Self-Refining Vision Language Model for Robotic Failure Detection and Reasoning

Carl Qi, Xiaojie Wang, Silong Yong +6

Reasoning about failures is crucial for building reliable and trustworthy robotic systems. Prior approaches either treat failure reasoning as a closed-set classification problem or…

cs.AI2026

Correct Reasoning Paths Visit Shared Decision Pivots

Dongkyu Cho, Amy B. Z. Zhang, Bilel Fehri +4

Chain-of-thought (CoT) reasoning exposes the intermediate thinking process of large language models (LLMs), yet verifying those traces at scale remains unsolved. In response, we in…

cs.LG2026

Hierarchical Entity-centric Reinforcement Learning with Factored Subgoal Diffusion

Dan Haramati, Carl Qi, Tal Daniel +3

We propose a hierarchical entity-centric framework for offline Goal-Conditioned Reinforcement Learning (GCRL) that combines subgoal decomposition with factored structure to solve l…

cs.AI2025

EC-Diffuser: Multi-Object Manipulation via Entity-Centric Behavior Generation

Carl Qi, Dan Haramati, Tal Daniel +2

Object manipulation is a common component of everyday tasks, but learning to manipulate objects from high-dimensional observations presents significant challenges. These challenges…