Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
LemonHarness Technical Report
Kailong Ren, Fubo Sun, Jiachen Liu +18
As large language model (LLM) agents are applied to longer tasks, they increasingly modify workspace state across multiple rounds of iteration. However, agents typically observe on…
cs.AI2026
From Brewing to Resolution: Tracing the Internal Lifecycle of Code Reasoning in LLMs
Siyue Chen, Yifu Guo, Yuquan Lu +9
Standard accuracy metrics cannot explain why LLMs handle variable tracking but fail on semantically equivalent loops. We study an internal lifecycle of code reasoning in which mode…