Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents
Mingguang Chen, Licheng Wang, Bo Qu
Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earl…
cs.CL2026
The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale
Mingguang Chen, Bo Qu, Licheng Wang
Accuracy changes after language-model self-revision are usually interpreted as changes in reasoning. We show this can fail at the answer-extraction boundary, and test the failure c…