8 papers
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
Boyang Zhang, Xiaobing Chen, Songyang Zhang +4
Mixture-of-Experts (MoE) models enable scalable neural networks through conditional computation, offering enhanced effectiveness and efficiency for next-generation wireless communi…
Towards Efficient Federated Learning of Networked Mixture-of-Experts for Mobile Edge Computing
Song Gao, Songyang Zhang, Shusen Jing +4
Recent advancements in large artificial intelligence models (LAMs) are driving significant innovations in mobile edge computing within next-generation wireless networks. However, t…
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
Mugunthan Shandirasegaran, Hongkang Li, Songyang Zhang +2
The recent empirical success of Mamba and other selective state space models (SSMs) has renewed interest in non-attention architectures for sequence modeling, yet their theoretical…
Theoretical Analysis of Contrastive Learning under Imbalanced Data: From Training Dynamics to a Pruning Solution
Haixu Liao, Yating Zhou, Songyang Zhang +2
Contrastive learning has emerged as a powerful framework for learning generalizable representations, yet its theoretical understanding remains limited, particularly under imbalance…
InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery
Shiyang Feng, Runmin Ma, Xiangchao Yan +53
We introduce InternAgent-1.5, a unified system designed for end-to-end scientific discovery across computational and empirical domains. The system is built on a structured architec…
Efficient Training of Large-Scale AI Models Through Federated Mixture-of-Experts: A System-Level Approach
Xiaobing Chen, Boyang Zhang, Xiangwei Zhou +4
The integration of Federated Learning (FL) and Mixture-of-Experts (MoE) presents a compelling pathway for training more powerful, large-scale artificial intelligence models (LAMs)…