2 papers
cs.AI2026
Process-Reward Tactic Evolution for Long-Horizon Bioinformatics Workflows
Lingzhi Yang, Yubo Fan, Song Wu +1
LLM agents can write code and call tools, but reliable bioinformatics work requires long-horizon interaction with workflow software, typed data objects, provenance, and biological…
cs.LG2026
STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
Haipeng Luo, Qingfeng Sun, Songli Wu +4
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from poli…