3 papers
q-bio.GN2026
Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min--Max Selection
Yaodi Luo, Peize He, Lingbei Meng +5
Large single-cell datasets are expensive to store, curate, and repeatedly reuse for model training. Data distillation can reduce this burden by building smaller training sets. Howe…
cs.CL2025
Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion
Jianqing Zhu, Huang Huang, Zhihang Lin +18
This paper addresses the critical need for democratizing large language models (LLM) in the Arab world, a region that has seen slower progress in developing models comparable to st…
cs.CL2024
Alignment at Pre-training! Towards Native Alignment for Arabic LLMs
Juhao Liang, Zhenyang Cai, Jianqing Zhu +9
The alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction…