collaborators

5 papers

cs.CL2025

Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models

Chen Yang, Guangyue Peng, Jiaying Zhu +16

We present Nanbeige4-3B, a family of small-scale but high-performing language models. Pretrained on 23T high-quality tokens and finetuned on over 30 million diverse instructions, w…

stat.ML2025

Non-stationary Bandit Convex Optimization: A Comprehensive Study

Xiaoqi Liu, Dorian Baudry, Julian Zimmert +2

Bandit Convex Optimization is a fundamental class of sequential decision-making problems, where the learner selects actions from a continuous domain and observes a loss (but not it…

cs.CL2025

Agentic Reinforcement Learning with Implicit Step Rewards

Xiaoqian Liu, Ke Wang, Yuchuan Wu +4

Large language models (LLMs) are increasingly developed as autonomous agents using reinforcement learning (agentic RL) that reason and act in interactive environments. However, spa…

cs.CL2025

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning

Xiaoqian Liu, Ke Wang, Yongbin Li +6

Large Language Models (LLMs) have shown impressive reasoning capabilities in well-defined problems with clear solutions, such as mathematics and coding. However, they still struggl…

cs.AI2025

SDPO: Segment-Level Direct Preference Optimization for Social Agents

Aobo Kong, Wentao Ma, Shiwan Zhao +7

Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex social dialogues. Direct Preference Optimization (DPO)…