2 papers
cs.CL2026
DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models
Yixin Bu, Runze Xia, Guanyun Zou +3
Accurate Uncertainty Quantification (UQ) is critical for reliable deployment of Large Language Models (LLMs), yet traditional probability-based metrics often fail to capture the mo…
cs.CL2026
Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents
Bin Wu, Guanyun Zou, Bingbing Wang +2
A long-lived LLM agent, such as OpenClaw, earns its value by acting on a user's preferences and constraints across sessions, not just the current request. Yet today's agents keep w…