3 papers
cs.AI2026
Mask-Proof: An LLM-based Automated Data Curation Pipeline on Mathematical Proofs
Jierui Zhang, Siyuan Tan, Xinhang Li +8
Large language models (LLMs) are increasingly capable of mathematical problem solving and can even assist with research-level proofs, yet we still lack a scalable and reproducible…
cs.LG2026
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
Wuyue Zhang, Chongdong Huang, Chunbo You +3
Training large-scale Mixture-of-Experts (MoE) models is bottlenecked by activation memory and expert-parallel communication, yet FP4 training remains impractical on Hopper-class GP…
cs.CL2026
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
Chenghao Fan, Zhenyi Lu, Sichen Liu +4
While Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning for Large Language Models (LLMs), its performance often falls short of Full Fine-Tuning (Full FT). Current…