3 papers
cs.LG2026
GRPO-QM: Target Preserving Exploration for Quantum Tomography
Yufeng Wang, Parivesh Priye, Lu Wei +1
Reward-based learning can alter the very posterior distribution that scientific inference aims to estimate. GRPO-QM sidesteps this by learning only an exploration strategy for a st…
cs.AI2026
VoiceLongMemEval: Do Assistants Remember How You Sounded?
Ramit Pahwa, Parivesh Priye, Apoorva Beedu
With the growing scale of multi-agent architectures and large language models, deployed AI assistants are increasingly tasked with reasoning over long, continuous, multi-session co…
cs.SD2026
Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
Ramit Pahwa, Apoorva Beedu, Parivesh Priye +4
Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acousti…