8 papers
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
Mechanistic Attention Guidance for Agent Memory Refinement
Yechao Hong, Haiquan Qiu, Yaqing Wang +1
Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections. However, this text-based paradigm rarely inco…
Graph Unitary Message Passing
Haiquan Qiu, Quanming Yao
Unitarity is a useful principle for stabilizing deep neural networks, but in graph neural networks (GNNs) instability is induced not only by learnable parameters but also by the gr…
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Haiquan Qiu, Quanming Yao
The pursuit of computational efficiency has driven the adoption of low-precision formats for training transformer models. However, this progress is often hindered by notorious trai…
Scaling GraphLLM with Bilevel-Optimized Sparse Querying
Yangzhe Peng, Haiquan Qiu, Quanming Yao +1
LLMs have recently shown strong potential in enhancing node-level tasks on text-attributed graphs (TAGs) by providing explanation features. However, their practical use is severely…
Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
Haiquan Qiu, You Wu, Yingjie Tan +2
Loss explosions in training deep neural networks can nullify multi-million dollar training runs. Conventional monitoring metrics like weight and gradient norms are often lagging an…