2 papers
cs.CL2026
ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents
Zheng Liu, Longxiang Zhang, Xintong Wang +8
LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on outcome-homogeneous groups wh…
cs.AI2026
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
Songlin Bai, Xintong Wang, Linlin Yu +12
In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every parameter must respect a regula…