Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach
Kihyun Kim, Shripad Deshmukh, Nikos Vlassis +1
Inverse reinforcement learning (IRL) typically assumes demonstrations from a single optimal demonstrator, but in many applications data come from multiple imperfect demonstrators w…
cs.LG2026
COSAC: Counterfactual Credit Assignment in Sequential Cooperative Teams
Shripad Deshmukh, Jayakumar Subramanian, Raghavendra Addanki +1
In cooperative teams where agents act in a fixed order and share a single team-level reward (multi-agent language systems, sequential robotic tasks), per-agent credit assignment is…