most citedConfigurable Foundation Models: Building LLMs from a Modular Perspective

4 citations · 5 across the 3 of their papers we have counts for

collaborators

5 papers

cs.LG2025

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity

Chenyang Song, Weilin Zhao, Xu Han +5

To alleviate the computational burden of large language models (LLMs), architectures with activation sparsity, represented by mixture-of-experts (MoE), have attracted increasing at…

cs.CL2025

Cost-Optimal Grouped-Query Attention for Long-Context Modeling

Yingfa Chen, Yutong Wu, Chenyang Song +5

Grouped-Query Attention (GQA) is a widely adopted strategy for reducing the computational cost of attention layers in large language models (LLMs). However, current GQA configurati…

cs.LG2024

Sparsing Law: Towards Large Language Models with Greater Activation Sparsity

Yuqi Luo, Chenyang Song, Xu Han +7

Activation sparsity denotes the existence of substantial weakly-contributed elements within activation outputs that can be eliminated, benefiting many important applications concer…

cs.AI20244 cited

Configurable Foundation Models: Building LLMs from a Modular Perspective

Chaojun Xiao, Zhengyan Zhang, Chenyang Song +20

Advancements in LLMs have recently unveiled challenges tied to computational efficiency and continual scalability due to their requirements of huge parameters, making the applicati…

cs.CL20241 cited

Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slip Scripts

Yingfa Chen, Chenlong Hu, Cong Feng +5

This study presents a multi-modal multi-granularity tokenizer specifically designed for analyzing ancient Chinese scripts, focusing on the Chu bamboo slip (CBS) script used during…