Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
Shuo He, Lang Feng, Qi Wei +3
Group-based reinforcement learning (RL), such as GRPO, has advanced the capabilities of large language models on long-horizon agentic tasks. To enable more fine-grained policy upda…
cs.LG2023
Regression with Cost-based Rejection
Xin Cheng, Yuzhou Cao, Haobo Wang +3
Learning with rejection is an important framework that can refrain from making predictions to avoid critical mispredictions by balancing between prediction and rejection. Previous…
cs.LG2023
Weakly Supervised Regression with Interval Targets
Xin Cheng, Yuzhou Cao, Ximing Li +2
This paper investigates an interesting weakly supervised regression setting called regression with interval targets (RIT). Although some of the previous methods on relevant regress…