activity
20212026
most citedMengzi: Towards Lightweight yet Ingenious Pre-trained Models for Chinese

30 citations · 33 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory

Hanzuo Liu, Xuan Qi, Chunyu Liu +6

Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for prediction. We turn this divisi…

cs.CL2024

Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities

Tianjie Ju, Yiting Wang, Xinbei Ma +7

The rapid adoption of large language models (LLMs) in multi-agent systems has highlighted their impressive capabilities in various applications, such as collaborative problem-solvi…

cs.CL2024★ 2 cited

Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models

Ang Lv, Yuhan Chen, Kaiyi Zhang +5

In this paper, we delve into several mechanisms employed by Transformer-based language models (LLMs) for factual recall tasks. We outline a pipeline consisting of three major steps…

cs.CL2024

On the Robustness of Editing Large Language Models

Xinbei Ma, Tianjie Ju, Jiyang Qiu +4

Large language models (LLMs) have played a pivotal role in building communicative AI, yet they encounter the challenge of efficient updates. Model editing enables the manipulation…

cs.CL2021★ 30 cited

Mengzi: Towards Lightweight yet Ingenious Pre-trained Models for Chinese

Zhuosheng Zhang, Hanqing Zhang, Keming Chen +4

Although pre-trained models (PLMs) have achieved remarkable improvements in a wide range of NLP tasks, they are expensive in terms of time and resources. This calls for the study o…