6 papers
A First-Principles Theory of Slow Thinking and Active Perception
Hongkang Yang, Zhi-Qin John Xu, Feiyu Xiong +1
As part of a series on first-principles modeling of cognitive functions, this paper attempts to provide a mathematical formulation of thinking and perception. It formally derives s…
Small Initialization Matters for Large Language Models
Liangkai Hang, Junjie Yao, Zhiyu Li +3
Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered. Although progress is usually attributed to…
Adaptive Preconditioners Trigger Loss Spikes in Adam
Zhiwei Bai, Zhangchen Zhou, Jiajie Zhao +6
Loss spikes commonly emerge during neural network training with the Adam optimizer across diverse architectures and scales, yet their underlying mechanism remains elusive. While pr…
MemOS: A Memory OS for AI System
Zhiyu Li, Chenyang Xi, Chunyu Li +36
Large Language Models (LLMs) have become an essential infrastructure for Artificial General Intelligence (AGI), yet their lack of well-defined memory management systems hinders the…
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
Liangkai Hang, Junjie Yao, Zhiwei Bai +17
The reasoning ability of large language models (LLMs) has been rapidly advancing in recent years, attracting interest in more fundamental approaches that can reliably enhance their…
MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
Zhiyu Li, Shichao Song, Hanyu Wang +19
Large Language Models (LLMs) have emerged as foundational infrastructure in the pursuit of Artificial General Intelligence (AGI). Despite their remarkable capabilities in language…