3 papers
cs.CL2026
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
cs.CV2026
PointCNN++: Performant Convolution on Native Points
Lihan Li, Haofeng Zhong, Rui Bu +4
Existing convolutional learning methods for 3D point cloud data are divided into two paradigms: point-based methods that preserve geometric precision but often face performance cha…
cs.LG2026
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
Rui Bu, Haofeng Zhong, Wenzheng Chen +1
Large models based on the Transformer architecture are susceptible to extreme-token phenomena, such as attention sinks and value-state drains. These issues, which degrade model per…