Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
ZenGen: Social Mind for LLMs
ZenGen Team, Zing Team, Ao Xiang +57
As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track…
cs.CL2024
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
Hongyu Wang, Shuming Ma, Ruiping Wang +1
We introduce, Q-Sparse, a simple yet effective approach to training sparsely-activated large language models (LLMs). Q-Sparse enables full sparsity of activations in LLMs which can…
cs.CL2024
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Shuming Ma, Hongyu Wang, Lingxiao Ma +7
Recent research, such as BitNet, is paving the way for a new era of 1-bit Large Language Models (LLMs). In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in whic…