10 papers
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
Dual-Process Atomic Skill Learning: Decoupling Semantic Reasoning and Real-Time Control
Jun Chen, Erdent Bao, Wenlong Dong +7
Language-conditioned Imitation Learning (IL) is essential for enabling robots to perform complex tasks following natural language instructions. However, generalizing to multi-step…
NFTR: From Provable Mode-Averaging to Geodesic Subgoal Selection in Offline Goal-Conditioned RL
Erdemt Bao, Xing Lei, Jun Chen
Hierarchical Implicit Q-Learning (HIQL), an offline goal-conditioned RL method, selects subgoals by value-function advantages alone. This rule has two coupled failure modes. Optimi…
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models
Jiyang Guan, Yong Xie, Jun Chen +6
Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains…
ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning
Chengcao Yang
We propose a paradigm shift toward open-ended curriculum self-play: rather than learning to answer on a fixed prompt set, a unified policy learns to question: generating verifiable…
Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning
Haoyu Wang, Haonan Wang, Yuyan Chen +5
In-context learning (ICL) allows large models to adapt to tasks using a few examples, yet its extension to vision-language models (VLMs) remains fragile. Our analysis reveals that…