4 papers · 1 filter
Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs
Lei Yang, Wei Bi, Chenxi Sun +2
On-policy reinforcement learning (RL) for language model post-training suffers from a fundamental tension: as training progresses, policy entropy collapses and sampling diversity d…
Knowledge Verification to Nip Hallucination in the Bud
Fanqi Wan, Xinting Huang, Leyang Cui +3
While large language models (LLMs) have demonstrated exceptional performance across various tasks following human alignment, they may still generate responses that sound plausible…
FuseChat: Knowledge Fusion of Chat Models
Fanqi Wan, Longguang Zhong, Ziyi Yang +2
While training large language models (LLMs) from scratch can indeed lead to models with distinct capabilities and strengths, it incurs substantial costs and may lead to redundancy…
Knowledge Fusion of Chat LLMs: A Preliminary Technical Report
Fanqi Wan, Ziyi Yang, Longguang Zhong +3
Recently, FuseLLM introduced the concept of knowledge fusion to transfer the collective knowledge of multiple structurally varied LLMs into a target LLM through lightweight continu…