5 citations · 12 across the 5 of their papers we have counts for
5 papers
Exploring the Benefit of Activation Sparsity in Pre-training
Zhengyan Zhang, Chaojun Xiao, Qiujieli Qin +7
Pre-trained Transformers inherently possess the characteristic of sparse activation, where only a small fraction of the neurons are activated for each token. While sparse activatio…
From MOOC to MAIC: Reshaping Online Teaching and Learning through LLM-driven Agents
Jifan Yu, Zheyuan Zhang, Daniel Zhang-li +30
Since the first instances of online education, where courses were uploaded to accessible and shared online platforms, this form of scaling the dissemination of human knowledge to r…
Configurable Foundation Models: Building LLMs from a Modular Perspective
Chaojun Xiao, Zhengyan Zhang, Chenyang Song +20
Advancements in LLMs have recently unveiled challenges tied to computational efficiency and continual scalability due to their requirements of huge parameters, making the applicati…
Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slip Scripts
Yingfa Chen, Chenlong Hu, Cong Feng +5
This study presents a multi-modal multi-granularity tokenizer specifically designed for analyzing ancient Chinese scripts, focusing on the Chu bamboo slip (CBS) script used during…
Iterative Experience Refinement of Software-Developing Agents
Chen Qian, Jiahao Li, Yufan Dang +8
Autonomous agents powered by large language models (LLMs) show significant potential for achieving high autonomy in various scenarios such as software development. Recent research…