1 citations · 1 across the 4 of their papers we have counts for
4 papers
Your Group-Relative Advantage Is Biased
Fengkai Yang, Zherui Chen, Xiaohan Wang +10
Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such…
LLMBoost: Make Large Language Models Stronger with Boosting
Zehao Chen, Tianxiang Ai, Yifei Li +11
Ensemble learning of LLMs has emerged as a promising alternative to enhance performance, but existing approaches typically treat models as black boxes, combining the inputs or fina…
FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits
Pingchen Lu, Zhi Hong, Zhiwei Shang +6
The performance of large language models (LLMs) is highly sensitive to the input prompt, making prompt optimization a critical task. However, real-world application is hindered by…
Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
Hongyan Xie, Yitong Yao, Yikun Ban +6
Large language models (LLMs) excel at reasoning tasks but are expensive to deploy. Thus small language models (SLMs) are fine-tuned on CoT data generated by LLMs to copy LLMs' abil…