5 papers
RH-Detect: A Unified Benchmark for Reward Hacking Detection
Junwei Quan, Evgenii Opryshko, Rohan Subramani +1
Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems. Existing datasets us…
World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
Junwei Quan, Evgenii Opryshko, Nicholas Rhinehart +1
Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and…
Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents
Sarthak Choudhary, Mihai Christodorescu, Ashish Hooda +3
Computer-use agents (CUAs) perform tasks across applications (such as desktops, mobile apps, and web browsers) by observing graphical interfaces and issuing commands such as clicks…
Test-Time Graph Search for Goal-Conditioned Reinforcement Learning
Evgenii Opryshko, Junwei Quan, Claas Voelcker +2
Offline goal-conditioned reinforcement learning (GCRL) often struggles with long-horizon tasks, where errors in value estimation accumulate and produce unreliable policies. It is t…
LEGOS-SLEEC: Tool for Formalizing and Analyzing Normative Requirements
Kevin Kolyakov, Lina Marsso, Nick Feng +2
Systems interacting with humans, such as assistive robots or chatbots, are increasingly integrated into our society. To prevent these systems from causing social, legal, ethical, e…