2 papers
cs.LG2026
VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning
Pengcheng Li, Zhengyang Zhang, Dongxu Zhang +2
Fine-grained credit assignment is a central challenge in reinforcement learning for long horizon LLM agents. Standard objectives often train from programmatically verifiable termin…
cs.AI2026
Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning
Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao +4
Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoni…