collaborators

7 papers

cs.LG2026

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Yan Wang, Qifan Zhang, Jiachen Yu +12

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead…

cs.CL2026

Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning

Xuewei Yang, Jiachen Yu, Jie Wu +3

Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in which increasingly concentrated…

cs.CL2026

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance

Jiachen Yu, Zhihao Xu, Junjie Wang +1

Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. Howeve…

cs.CL2026

CoWork-X: Experience-Optimized Co-Evolution for Multi-Agent Collaboration System

Zexin Lin, Jiachen Yu, Haoyang Zhang +5

Large language models are enabling language-conditioned agents in interactive environments, but highly cooperative tasks often impose two simultaneous constraints: sub-second real-…

cs.CL2025

S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models

Shaoning Sun, Jiachen Yu, Zongqi Wang +3

With the rapid development of large language models (LLMs), generative reward models (GRMs) have been widely adopted for reward modeling and evaluation. Previous studies have prima…

cs.CL2025

Improve LLM-as-a-Judge Ability as a General Ability

Jiachen Yu, Shaoning Sun, Xiaohui Hu +3

LLM-as-a-Judge leverages the generative and reasoning capabilities of large language models (LLMs) to evaluate LLM responses across diverse scenarios, providing accurate preference…