activity
20242026
most citedTrustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning

1 citations · 1 across the 6 of their papers we have counts for

collaborators

8 papers

cs.DC2026

DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale

Jialiang Huang, Hongxuan Tang, Jingchang Chen +128

Large-scale agentic training and evaluation with large language models (LLMs) rely on isolated, stateful execution environments in which models inspect repositories, invoke tools,…

cs.CL2026

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, :, Anyi Xu +585

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation…

cs.SE2026

CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation

Jianfeng Cai, Jinhua Zhu, Ruopei Sun +5

The rise of reasoning models necessitates large-scale verifiable data, for which programming tasks serve as an ideal source. However, while competitive programming platforms provid…

cs.AI2025

Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks

Ruopei Sun, Jianfeng Cai, Jinhua Zhu +5

RLHF has emerged as a predominant approach for aligning artificial intelligence systems with human preferences, demonstrating exceptional and measurable efficacy in instruction fol…

cs.LG2025

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

Kangwen Zhao, Jianfeng Cai, Jinhua Zhu +5

Reinforcement Learning from Human Feedback (RLHF) relies on reward models to align large language models with human preferences. However, RLHF often suffers from reward hacking, wh…

cs.LG2025

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling

Jianfeng Cai, Jinhua Zhu, Ruopei Sun +4

Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward m…