works on

From the 1 of 11 linked papers with an AI index.

collaborators

11 papers

cs.LG2026

Speculate with Memory: Lossless Acceleration for LLM Agents

Yu Li, Qinyuan Ye, Prafulla Kumar Choubey +2

The paper proposes adding online memory systems to speculative execution for large language model agents, enabling the speculator to learn from past trajectories and improve predic…

cs.LG2026

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

Yuanda Xu, Zhengze Zhou, Hejian Sang +6

Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO…

cs.CL2026

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion

Haoyi Qiu, Yilun Zhou, Pranav Narayanan Venkit +4

As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored the dynamics of Agent-to-Agent (A2…

cs.CL2026

Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination

Prafulla Kumar Choubey, Kung-Hsiang Huang, Pranav Narayanan Venkit +5

Enterprise deep research often fails to produce decision-ready reports due to uneven information coverage, context explosion, and premature stopping. We propose a scalable Enterpri…

cs.AI2026

From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models

Jiaxin Zhang, Wendi Cui, Zhuohang Li +4

While Large Language Models (LLMs) show remarkable capabilities, their unreliability remains a critical barrier to deployment in high-stakes domains. This survey charts a functiona…

cs.LG2026

The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation

Jiaxin Zhang, Xiangyu Peng, Qinglin Chen +3

On-policy distillation (OPD) is an increasingly important paradigm for post-training language models. However, we identify a pervasive Scaling Law of Miscalibration: while OPD effe…