collaborators

5 papers

cs.CL2026

Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems

Kartikey Singh Bhandari, Aarya Wadhwani, Dhruv Kumar +1

LLM agents that persist across sessions accumulate stored memories whose validity varies enormously by content type, yet existing memory architectures treat all memories as equally…

cs.AI2026

LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo

Ojas Jain, Dhruv Kumar

We introduce LudoBench, a benchmark for evaluating LLM strategic reasoning in Ludo, a stochastic multi-agent board game whose dice mechanics, piece capture, safe-square navigation,…

cs.CL2026

Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows

Aditya Sharan, Sriram Hebbale, Dhruv Kumar

Training large language models for complex reasoning is bottlenecked by the scarcity of verifiable, high-quality data. In domains like physics, standard text augmentation often int…

cs.LG2026

Generative Evolutionary Meta-Solver (GEMS): Scalable Surrogate-Free Multi-Agent Reinforcement Learning

Alakh Sharma, Gaurish Trivedi, Kartikey Singh Bhandari +4

Scalable multi-agent reinforcement learning (MARL) remains a central challenge for AI. Existing population-based methods, like Policy-Space Response Oracles, PSRO, require storing…

cs.MA2025

Multi-Agent Inverse Q-Learning from Demonstrations

Nathaniel Haynam, Adam Khoja, Dhruv Kumar +2

When reward functions are hand-designed, deep reinforcement learning algorithms often suffer from reward misspecification, causing them to learn suboptimal policies in terms of the…