2 papers
cs.CL2025
ToolExpander: Extending the Frontiers of Tool-Using Reinforcement Learning to Weak LLMs
Fu Chen, Peng Wang, Xiyin Li +3
Training Large Language Models (LLMs) with Group Relative Policy Optimization (GRPO) encounters a significant challenge: models often fail to produce accurate responses, particular…
cs.LG2021
Binary Classification from Multiple Unlabeled Datasets via Surrogate Set Classification
Nan Lu, Shida Lei, Gang Niu +2
To cope with high annotation costs, training a classifier only from weakly supervised data has attracted a great deal of attention these days. Among various approaches, strengtheni…