1 citations · 1 across the 1 of their papers we have counts for
10 papers
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
Xiaodong Lu, Xiaohan Wang, Jiajun Chai +7
Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods uti…
Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
Zhijun Chen, Xiaodong Lu, Jingzheng Li +12
LLM Ensemble -- which involves the comprehensive use of multiple large language models (LLMs), each aimed at handling user queries during downstream inference, to benefit from thei…
Silent Inconsistency in Data-Parallel Full Fine-Tuning: Diagnosing Worker-Level Optimization Misalignment
Hong Li, Zhen Zhou, Honggang Zhang +4
Data-parallel (DP) training with synchronous all-reduce is a dominant paradigm for full-parameter fine-tuning of large language models (LLMs). While parameter synchronization guara…
ChiEngMixBench: Evaluating Large Language Models on Expert-Style Chinese-English Terminology Mixing
Qingyan Yang, Tongxi Wang, Yunsheng Luo
Large language models increasingly mediate multilingual professional communication, where useful generation requires adapting to community conventions about which expressions are r…
Towards Robust Zero-Shot Reinforcement Learning
Kexin Zheng, Lauriane Teyssier, Yinan Zheng +2
The recent development of zero-shot reinforcement learning (RL) has opened a new avenue for learning pre-trained generalist policies that can adapt to arbitrary new tasks in a zero…
Towards Real-Time Fake News Detection under Evidence Scarcity
Guangyu Wei, Ke Han, Yueming Lyu +4
Fake news detection becomes particularly challenging in real-time scenarios, where emerging events often lack sufficient supporting evidence. Existing approaches often rely heavily…