collaborators

6 papers

cs.CL2026

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +398

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…

cs.LG2026

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

Kangwen Zhao, Jianfeng Cai, Jinhua Zhu +5

Reinforcement Learning from Human Feedback (RLHF) relies on reward models to align large language models with human preferences. However, RLHF often suffers from reward hacking, wh…

cs.SE2026

CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation

Jianfeng Cai, Jinhua Zhu, Ruopei Sun +5

The rise of reasoning models necessitates large-scale verifiable data, for which programming tasks serve as an ideal source. However, while competitive programming platforms provid…

cs.AI2025

Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks

Ruopei Sun, Jianfeng Cai, Jinhua Zhu +5

RLHF has emerged as a predominant approach for aligning artificial intelligence systems with human preferences, demonstrating exceptional and measurable efficacy in instruction fol…

cs.CV2025

Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering

Jianfeng Cai, Wengang Zhou, Zongmeng Zhang +3

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.However, hallucination, where the model generates plausible yet incorrect outputs,…

cs.LG2025

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling

Jianfeng Cai, Jinhua Zhu, Ruopei Sun +4

Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward m…