5 citations · 11 across the 17 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Dynamic Important Example Mining for Reinforcement Finetuning
Haoru Tan, Sitong Wu, Yanfeng Chen +9
Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and use…
cs.AI2024★ 5 cited
Advancing LLM Reasoning Generalists with Preference Trees
Lifan Yuan, Ganqu Cui, Hanbin Wang +12
We introduce Eurus, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B and CodeLlama-70B, Eurus models achieve state-of-the-art results amon…