3 papers
cs.LG2026
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
Tanmay Ambadkar, Sourav Panda, Shreyash Kale +2
Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalabl…
cs.AI2026
Two-Bridge: Exclusive Objectives and Extended Horizon StarCraft II Benchmark
Sourav Panda, Tanmay Ambadkar, Shreyash Kale +2
The research community lacks a middle ground between StarCraft II full game and its mini-games. The full-game's sprawling state-action space renders reward signals sparse and noisy…
cs.AI2026
Automating the Refinement of Reinforcement Learning Specifications
Tanmay Ambadkar, ÄorÄe ŽikeliÄ, Abhinav Verma
Logical specifications have been shown to help reinforcement learning algorithms in achieving complex tasks. However, when a task is under-specified, agents might fail to learn use…