works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CL2026

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +398

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…

cs.CV2026

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

Xiaokang Ma, Yifan Sun, Zhihong Jin +6

The paper introduces ReflectWorld-MM, an entity-oriented multimodal memory architecture that processes continuous video streams, stores observations in a hierarchical long‑term mem…

cs.CV2026

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective

Lei Lei, Jie Gu, Xiaokang Ma +3

Existing Multimodal Large Language Models (MLLMs) process a large number of visual tokens, leading to significant computational costs and inefficiency. Instruction-related visual t…

cs.RO2025

Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model

Bokai Ji, Jie Gu, Xiaokang Ma +3

Affordance is crucial for intelligent robots in the context of object manipulation. In this paper, we argue that affordance should be task-/instruction-dependent, which is overlook…

cs.RO2025

Stimulating Imagination: Towards General-purpose "Something Something Placement"

Jianyang Wu, Jie Gu, Xiaokang Ma +3

General-purpose object placement is a fundamental capability of an intelligent generalist robot: being capable of rearranging objects following precise human instructions even in n…

cs.AI2025

Kimi k1.5: Scaling Reinforcement Learning with LLMs

Kimi Team, Angang Du, Bofei Gao +93

Language model pretraining with next token prediction has proved effective for scaling compute but is limited to the amount of available training data. Scaling reinforcement learni…