most citedContextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

1 citations · 1 across the 1 of their papers we have counts for

collaborators

10 papers

cs.LG20261 cited

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

Xiaodong Lu, Xiaohan Wang, Jiajun Chai +7

Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods uti…

cs.CL2026

Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

Zhijun Chen, Xiaodong Lu, Jingzheng Li +12

LLM Ensemble -- which involves the comprehensive use of multiple large language models (LLMs), each aimed at handling user queries during downstream inference, to benefit from thei…

cs.LG2026

Silent Inconsistency in Data-Parallel Full Fine-Tuning: Diagnosing Worker-Level Optimization Misalignment

Hong Li, Zhen Zhou, Honggang Zhang +4

Data-parallel (DP) training with synchronous all-reduce is a dominant paradigm for full-parameter fine-tuning of large language models (LLMs). While parameter synchronization guara…

cs.CL2026

ChiEngMixBench: Evaluating Large Language Models on Expert-Style Chinese-English Terminology Mixing

Qingyan Yang, Tongxi Wang, Yunsheng Luo

Large language models increasingly mediate multilingual professional communication, where useful generation requires adapting to community conventions about which expressions are r…

cs.LG2025

Towards Robust Zero-Shot Reinforcement Learning

Kexin Zheng, Lauriane Teyssier, Yinan Zheng +2

The recent development of zero-shot reinforcement learning (RL) has opened a new avenue for learning pre-trained generalist policies that can adapt to arbitrary new tasks in a zero…

cs.CL2025

Towards Real-Time Fake News Detection under Evidence Scarcity

Guangyu Wei, Ke Han, Yueming Lyu +4

Fake news detection becomes particularly challenging in real-time scenarios, where emerging events often lack sufficient supporting evidence. Existing approaches often rely heavily…