5 papers
SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
Ziyi Yang, Weizhou Shen, Chenliang Li +5
Progress in long-context reasoning for large language models (LLMs) has lagged behind other recent advances. This gap arises not only from the intrinsic difficulty of processing lo…
Lookahead Routing for Large Language Models
Canbin Huang, Tianyuan Shi, Yuhua Zhu +2
Large language model (LLM) routers improve the efficiency of multi-model systems by directing each query to the most appropriate model while leveraging the diverse strengths of het…
Discriminative Policy Optimization for Token-Level Reward Models
Hongzhan Chen, Tao Yang, Shiping Gao +4
Process reward models (PRMs) provide more nuanced supervision compared to outcome reward models (ORMs) for optimizing policy models, positioning them as a promising approach to enh…
BlockPruner: Fine-grained Pruning for Large Language Models
Longguang Zhong, Fanqi Wan, Ruijun Chen +2
With the rapid growth in the size and complexity of large language models (LLMs), the costs associated with their training and inference have escalated significantly. Research indi…
Knowledge Distillation of Black-Box Large Language Models
Hongzhan Chen, Ruijun Chen, Yuqi Yi +4
Given the exceptional performance of proprietary large language models (LLMs) like GPT-4, recent research has increasingly focused on boosting the capabilities of smaller models th…