7 papers
Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
Varun Giridhar, Anant Khandelwal, Jeremy A. Collins +2
Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from…
Read the Trace, Steer the Path: Trajectory-Aware Reinforcement Learning for Diffusion Language Models
Anant Khandelwal, Manish Gupta
Diffusion large language models (dLLMs) generate responses by iteratively unmasking and revising many positions in parallel. This process leaves a rich denoising trace depicting wh…
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
Subba Reddy Oota, Anant Khandelwal, Khushbu Pahwa +4
Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience a…
How does longer temporal context enhance multimodal narrative video processing in the brain?
Prachi Jindal, Anant Khandelwal, Manish Gupta +3
Understanding how humans and artificial intelligence systems process complex narrative videos is a fundamental challenge at the intersection of neuroscience and machine learning. T…
Linguistic properties and model scale in brain encoding: from small to compressed language models
Subba Reddy Oota, Vijay Rowtula, Satya Sai Srinath Namburi +5
Recent work has shown that scaling large language models (LLMs) improves their alignment with human brain activity, yet it remains unclear what drives these gains and which represe…
HistoryBankQA: Multilingual Temporal Question Answering on Historical Events
Biswadip Mandal, Anant Khandelwal, Manish Gupta
Temporal reasoning about historical events is a critical skill for NLP tasks like event extraction, historical entity linking, temporal question answering, timeline summarization,…