collaborators

9 papers

cs.CL2026

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +398

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…

cs.CV2026

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

Mingqi Gao, Hongyuan Dong, Yifei Chen +6

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense vi…

cs.LG2026

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Wenjin Hou, Shangpin Peng, Weinong Wang +13

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model…

cs.IR2026

Graph-GRPO: Dependency-Aware Credit Assignment for Generative E-commerce Search Relevance

Jiarui Che, Yifei Chen, Zhixing Tian +2

Search relevance modeling is a core task in e-commerce search systems, assessing how well a user query matches candidate products. Rather than relying on a single holistic matching…

cs.AI2026

ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration

Yifei Chen, Guanting Dong, Zhicheng Dou

Large Language Models (LLMs) can extend their parameter knowledge limits by adopting the Tool-Integrated Reasoning (TIR) paradigm. However, existing LLM-based agent training framew…

cs.AI2025

Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning

Yifei Chen, Guanting Dong, Zhicheng Dou

Tool-Integrated Reasoning (TIR) enables large language models (LLMs) to improve their internal reasoning ability by integrating external tools. However, models employing TIR often…