21 papers
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring
Renbiao Jin, Mingxin Yang, Yutian Chen +8
Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet robust restoration is critica…
4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture
Yutian Chen, Shi Guo, Tianshuo Yang +4
Reconstructing fast-dynamic scenes from multi-view videos is crucial for high-speed motion analysis and realistic 4D reconstruction. However, the majority of 4D capture systems are…
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
Sirui Chen, Lei Xu, Yuying Zhao +6
Recent RL methods have substantially improved the reasoning abilities of LLMs. Existing reward designs mainly follow two paradigms: (1) Reinforcement learning with verifiable rewar…
Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers
Yutian Chen, Yuheng Qiu, Ruogu Li +4
We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the base model. Co-Me distilled a l…
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
Tianshuo Yang, Guanyu Chen, Yutian Chen +8
While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often compromises the profound rea…