Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations
Da Ma, Gonghu Shang, Zhi Chen +6
Instruction tuning improves the ability of large language models (LLMs) to follow diverse human instructions, but achieving strong performance on specific target tasks remains chal…
cs.LG2025
LESA: Learnable LLM Layer Scaling-Up
Yifei Yang, Zouying Cao, Xinbei Ma +4
Training Large Language Models (LLMs) from scratch requires immense computational resources, making it prohibitively expensive. Model scaling-up offers a promising solution by leve…