6 citations · 14 across the 14 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025★ 1 cited
Delta Decompression for MoE-based LLMs Compression
Hao Gu, Wei Li, Lujun Li +5
Mixture-of-Experts (MoE) architectures in large language models (LLMs) achieve exceptional performance, but face prohibitive storage and memory requirements. To address these chall…
cs.LG2024★ 2 cited
NoRA: Nested Low-Rank Adaptation for Efficient Fine-Tuning Large Models
Cheng Lin, Lujun Li, Dezhi Li +3
In this paper, we introduce Nested Low-Rank Adaptation (NoRA), a novel approach to parameter-efficient fine-tuning that extends the capabilities of Low-Rank Adaptation (LoRA) techn…
cs.LG2024
STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
Peijie Dong, Lujun Li, Yuedong Zhong +8
In this paper, we present the first structural binarization method for LLM compression to less than 1-bit precision. Although LLMs have achieved remarkable performance, their memor…