collaborators

5 papers

cs.LG2026

Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference

Baihui Liu, Kaiyuan Tian, Wei Wang +3

Mixture-of-Experts (MoE) has become a dominant architecture for scaling large language models due to their sparse activation mechanism. However, the substantial number of expert ac…

cs.CL2026

GRASS: Gradient-based Adaptive Layer-wise Importance Sampling for Memory-efficient Large Language Model Fine-tuning

Kaiyuan Tian, Yu Tang, Gongqingjian Jiang +5

Full-parameter fine-tuning of large language models is constrained by substantial GPU memory requirements. Low-rank adaptation methods mitigate this challenge by updating only a su…

cs.CL2025

Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs

Wanyang Hong, Zhaoning Zhang, Yi Chen +5

Large Language Models (LLMs) have achieved remarkable performance on single-turn tasks, yet their effectiveness deteriorates in multi-turn conversations. We define this phenomenon…

cs.LG2025

ParaDySe: A Parallel-Strategy Switching Framework for Dynamic Sequence Lengths in Transformer

Zhixin Ou, Peng Liang, Jianchen Han +2

Dynamic sequences with varying lengths have been widely used in the training of Transformer-based large language models (LLMs). However, current training frameworks adopt a pre-def…

cs.LG2025

A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science

Kaiyuan Tian, Linbo Qiao, Baihui Liu +3

Scientific research faces high costs and inefficiencies with traditional methods, but the rise of deep learning and large language models (LLMs) offers innovative solutions. This s…