activity
20242026
collaborators

5 papers

cs.LG2026

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization

Tanmay Ambadkar, Sourav Panda, Shreyash Kale +2

Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalabl…

cs.AI2026

Two-Bridge: Exclusive Objectives and Extended Horizon StarCraft II Benchmark

Sourav Panda, Tanmay Ambadkar, Shreyash Kale +2

The research community lacks a middle ground between StarCraft II full game and its mini-games. The full-game's sprawling state-action space renders reward signals sparse and noisy…

cs.CR2026

Attention Is Where You Attack

Aviral Srivastava, Sourav Panda

Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety behavior remain poorly understo…

cs.HC2025

Signed, Sealed,... Confused: Exploring the Understandability and Severity of Policy Documents

Shikha Soneji, Sourav Panda, Sameer Neve +1

In general, Terms of Service (ToS) and other policy documents are verbose and full of legal jargon, which poses challenges for users to understand. To improve user accessibility an…

cs.CR2024

A Formal Framework for Assessing and Mitigating Emergent Security Risks in Generative AI Models: Bridging Theory and Dynamic Risk Mitigation

Aviral Srivastava, Sourav Panda

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overl…