activity
20242026
collaborators

10 papers

cs.LG2026

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control

Rohit Kumar Salla, Manoj Saravanan, Simon Stepputtis

Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottleneck. Reducing this cost requ…

cs.RO2026

From Local Corrections to Generalized Skills: Improving Neuro-Symbolic Policies with MEMO

Benjamin A. Christie, Yinlong Dai, Mohammad Bararjanianbahnamiri +2

Recent works use a neuro-symbolic framework for general manipulation policies. The advantage of this framework is that -- by applying off-the-shelf vision and language models -- th…

cs.AI2026

Bayesian Social Deduction with Graph-Informed Language Models

Shahab Rahimirad, Guven Gergerli, Lucia Romero +4

Social reasoning - inferring unobservable beliefs and intentions from partial observations of other agents - remains a challenging task for large language models (LLMs). We evaluat…

cs.LG2026

Overcoming Valid Action Suppression in Unmasked Policy Gradient Algorithms

Renos Zabounidis, Roy Siegelmann, Mohamad Qadri +3

In reinforcement learning environments with state-dependent action validity, action masking consistently outperforms penalty-based handling of invalid actions, yet existing theory…

cs.LG2026

SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding

Renos Zabounidis, Yue Wu, Simon Stepputtis +4

LM-based agents excel when given high-level action APIs but struggle to ground language into low-level control. Prior work has LLMs generate skills or reward functions for RL, but…

cs.AI2026

LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models

Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +5

Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire inc…