3 papers
cs.AI2026
ContraPrompt: Contrastive Prompt Optimization via Dyadic Reasoning Trace Analysis
Rishav Rishav, Pushpak Pujari, Pushpendre Rastogi
Prompt optimization methods either analyze individual failures in isolation or compare prompt variants across examples, operating on single execution traces with no access to the r…
cs.AI2025
Behaviour Discovery and Attribution for Explainable Reinforcement Learning
Rishav Rishav, Somjit Nath, Vincent Michalski +1
Building trust in reinforcement learning (RL) agents requires understanding why they make certain decisions, especially in high-stakes applications like robotics, healthcare, and f…
cs.LG2025
Handling Delay in Real-Time Reinforcement Learning
Ivan Anokhin, Rishav Rishav, Matthew Riemer +3
Real-time reinforcement learning (RL) introduces several challenges. First, policies are constrained to a fixed number of actions per second due to hardware limitations. Second, th…