From the 1 of 10 linked papers with an AI index.
10 papers
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
Yutong He, Daibo Li, Guohong Li +15
ReasFlow is an autonomous multi‑agent system that leverages large language models to perform rigorous mathematical reasoning, retrieve relevant knowledge, and generate complete res…
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
Boao Kong, Weichen Jia, Engao Zhang +6
Training stability is a key bottleneck in low-precision language model training: efficient low-cost paths can still produce short-lived numerical risks at a small set of operators.…
Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning
Bing Liu, Boao Kong, Limin Lu +2
Decentralized learning often involves a weighted global loss with heterogeneous node weights . We revisit two natural strategies for incorporating these weights: (i) embedding t…
CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
Boao Kong, Junzhu Liang, Yuxi Liu +2
Low-rank architectures have become increasingly important for efficient large language model (LLM) pre-training, providing substantial reductions in both parameter complexity and m…
BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization
Hengrui Zhang, Boao Kong, Engao Zhang +1
Stochastic bilevel optimization (SBO) has become a standard framework for hyperparameter learning, data reweighting, representation learning, and data-mixture optimization in deep…
SUDA-Muon: Structural Design Principles and Boundaries for Fully Decentralized Muon
Hengrui Zhang, Boao Kong, Jiahe Geng +1
Fully decentralized Muon is difficult because its nonlinear matrix-sign operator does not commute with linear gossip averaging. This makes decentralized Muon a structural design pr…