7 papers
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
Guangyu Zhao, Kewei Lian, Haoxuan Ru +8
Goal-conditioned policies enable decision-making models to execute diverse behaviors based on specified goals, yet their downstream performance is often highly sensitive to the cho…
Deep (Predictive) Discounted Counterfactual Regret Minimization
Hang Xu, Kai Li, Haobo Fu +3
Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. To enhance CFR's applicability in large games, researchers u…
Hamiltonian Descent Algorithms for Optimization: Accelerated Rates via Randomized Integration Time
Qiang Fu, Andre Wibisono
We study the Hamiltonian flow for optimization (HF-opt), which simulates the Hamiltonian dynamics for some integration time and resets the velocity to to decrease the objective…
Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning
Jinmin He, Kai Li, Yifan Zang +4
Offline multi-task reinforcement learning aims to learn a unified policy capable of solving multiple tasks using only pre-collected task-mixed datasets, without requiring any onlin…
Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance
Jinmin He, Kai Li, Yifan Zang +4
Multi-task reinforcement learning endeavors to efficiently leverage shared information across various tasks, facilitating the simultaneous learning of multiple tasks. Existing appr…