◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Aditya Modi

4 papers hereh-index 214 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3
  • last author1

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.AI1

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2026

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +4

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot gene…

cs.AI2026

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

Ankur Samanta, Akshayaa Magesh, Tal Lancewicki +7

Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environm…

cs.LG2026

Formalizing Learning from Language Feedback with Provable Guarantees

Wanqiao Xu, Allen Nie, Ruijie Zheng +3

Interactively learning from observation and language feedback is an increasingly studied area driven by the emergence of large language model (LLM) agents. Despite impressive empir…

cs.LG2025

How to Solve Contextual Goal-Oriented Problems with Offline Datasets?

Ying Fan, Jingling Li, Adith Swaminathan +2

We present a novel method, Contextual goal-Oriented Data Augmentation (CODA), which uses commonly available unlabeled trajectories and context-goal pairs to solve Contextual Goal-O…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.