Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
Chen Chen, Lai Wei
Large language model (LLM) scaling is hitting a wall. Widening models yields diminishing returns, and extending context length does not improve fundamental expressivity. In contras…
cs.LG2026
Monkey Jump : MoE-Style PEFT for Efficient Multi-Task Learning
Nusrat Jahan Prottasha, Md Kowsher, Chun-Nam Yu +2
Mixture-of-experts variants of parameter-efficient fine-tuning enable per-token specialization, but they introduce additional trainable routers and expert parameters, increasing me…