6 papers
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
Kangwen Zhao, Jianfeng Cai, Jinhua Zhu +5
Reinforcement Learning from Human Feedback (RLHF) relies on reward models to align large language models with human preferences. However, RLHF often suffers from reward hacking, wh…
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
Jianfeng Cai, Jinhua Zhu, Ruopei Sun +5
The rise of reasoning models necessitates large-scale verifiable data, for which programming tasks serve as an ideal source. However, while competitive programming platforms provid…
Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
Ruopei Sun, Jianfeng Cai, Jinhua Zhu +5
RLHF has emerged as a predominant approach for aligning artificial intelligence systems with human preferences, demonstrating exceptional and measurable efficacy in instruction fol…
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
Jianfeng Cai, Wengang Zhou, Zongmeng Zhang +3
Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.However, hallucination, where the model generates plausible yet incorrect outputs,…
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling
Jianfeng Cai, Jinhua Zhu, Ruopei Sun +4
Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward m…