2 papers
cs.LG2026
Mitigating False Credit Propagation: Probabilistic Graphical Reward Aggregation for Rubric-Based Reinforcement Learning
Can Lv, Mingju Chen, Heng Chang +1
Rubric-based rewards are increasingly used for open-ended language model post-training, but criterion-level scores are often aggregated as independent utilities. This flat scalariz…
cs.CL2026
HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
Mingju Chen, Can Lv, Guibin Zhang +2
LLM agents are increasingly expected to operate across heterogeneous task regimes that require distinct execution paradigms. This challenges fixed agent systems and motivates syste…