1 citations · 1 across the 6 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
Zhijun Tu, Jian Li, Yuanyuan Xi +5
1-bit LLM quantization offers significant advantages in reducing storage and computational costs. However, existing methods typically train 1-bit LLMs from scratch, failing to full…
cs.CL2024
MooER: LLM-based Speech Recognition and Translation Models from Moore Threads
Junhao Xu, Zhenlin Liang, Yi Liu +5
In this paper, we present MooER, a LLM-based large-scale automatic speech recognition (ASR) / automatic speech translation (AST) model of Moore Threads. A 5000h pseudo labeled data…
cs.CL2024★ 1 cited
Aurora:Activating Chinese chat capability for Mixtral-8x7B sparse Mixture-of-Experts through Instruction-Tuning
Rongsheng Wang, Haoming Chen, Ruizhe Zhou +8
Existing research has demonstrated that refining large language models (LLMs) through the utilization of machine-generated instruction-following data empowers these models to exhib…