6 papers
Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems
Kartikey Singh Bhandari, Aarya Wadhwani, Dhruv Kumar +1
LLM agents that persist across sessions accumulate stored memories whose validity varies enormously by content type, yet existing memory architectures treat all memories as equally…
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
Kartikey Singh Bhandari, Tanish Jain, Archit Agrawal +3
Customer reviews contain valuable signals about service quality, but converting large-scale review corpora into actionable business recommendations remains difficult. Standard sent…
Generative Evolutionary Meta-Solver (GEMS): Scalable Surrogate-Free Multi-Agent Reinforcement Learning
Alakh Sharma, Gaurish Trivedi, Kartikey Singh Bhandari +4
Scalable multi-agent reinforcement learning (MARL) remains a central challenge for AI. Existing population-based methods, like Policy-Space Response Oracles, PSRO, require storing…
Trust Regions Sell, But Who's Buying? Overlap Geometry as an Alternative Trust Region for Policy Optimization
Gaurish Trivedi, Alakh Sharma, Kartikey Singh Bhandari +4
Standard trust-region methods constrain policy updates via Kullback-Leibler (KL) divergence. However, KL controls only an average divergence and does not directly prevent rare, lar…
Actionable Advice from Reviews via Mixture of LoRA Experts: A Two-LLM Pipeline for Issue Extraction and Business Recommendations
Kartikey Singh Bhandari, Manav Ganesh, Yashwant Viswanathan +3
Customer reviews contain detailed, domain specific signals about service failures and user expectations, but converting this unstructured feedback into actionable business decision…
HAEPO: History-Aggregated Exploratory Policy Optimization
Gaurish Trivedi, Alakh Sharma, Kartikey Singh Bhandari +3
Exploration is essential in modern learning, from reinforcement learning environments with small neural policies to large language models (LLMs). Existing work, such as DPO, levera…