5 papers
ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko +1
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended re…
Modification-Considering Value Learning for Reward Hacking Mitigation in RL
Evgenii Opryshko, Umangi Jain, Igor Gilitschenski
Reinforcement learning agents can exploit misspecified reward signals to achieve high apparent returns while failing on the intended objective, a failure mode known as reward hacki…
Update-Free On-Policy Steering via Verifiers
Maria Attarian, Ian Vyse, Claas Voelcker +7
In recent years, Behavior Cloning (BC) has become one of the most prevalent methods for learning manipulation from human demonstrations. Despite their successes, BC policies are of…
Test-Time Graph Search for Goal-Conditioned Reinforcement Learning
Evgenii Opryshko, Junwei Quan, Claas Voelcker +2
Offline goal-conditioned reinforcement learning (GCRL) often struggles with long-horizon tasks, where errors in value estimation accumulate and produce unreliable policies. It is t…
Robust Reasoning Benchmark
Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko +1
While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on the context and textual formatting. We in…