2 papers
cs.LG2026
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
Tanmay Ambadkar, Sourav Panda, Shreyash Kale +2
Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalabl…
cs.AI2026
Two-Bridge: Exclusive Objectives and Extended Horizon StarCraft II Benchmark
Sourav Panda, Tanmay Ambadkar, Shreyash Kale +2
The research community lacks a middle ground between StarCraft II full game and its mini-games. The full-game's sprawling state-action space renders reward signals sparse and noisy…