3 papers
cs.AI2026
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
Chinmaya Kausik, Adith Swaminathan, Nathan Kallus
Large Language Model (LLM) agents are deployed in complex environments -- such as massive codebases, enterprise databases, and conversational histories -- where the relevant state…
cs.LG2025
Leveraging Offline Data in Linear Latent Contextual Bandits
Chinmaya Kausik, Kevin Tan, Ambuj Tewari
Leveraging offline data is an attractive way to accelerate online sequential decision-making. However, it is crucial to account for latent states in users or environments in the of…
cs.LG2024
A Theoretical Framework for Partially Observed Reward-States in RLHF
Chinmaya Kausik, Mirco Mutti, Aldo Pacchiano +1
The growing deployment of reinforcement learning from human feedback (RLHF) calls for a deeper theoretical investigation of its underlying models. The prevalent models of RLHF do n…