3 papers
cs.LG2026
Confidence-Adaptive SwiGLU for Mixture-of-Experts
Shaohua Li, Xiuchao Sui, Xiaobing Sun +4
SwiGLU has become a standard gated activation in modern Transformer MLPs, yet its gate sharpness -- the smoothness and selectivity of the gating function -- is typically fixed thro…
cs.LG2025
Rethinking Hebbian Principle: Low-Dimensional Structural Projection for Unsupervised Learning
Shikuang Deng, Jiayuan Zhang, Yuhang Wu +2
Hebbian learning is a biological principle that intuitively describes how neurons adapt their connections through repeated stimuli. However, when applied to machine learning, it su…
cs.LG2025
Temporal Flexibility in Spiking Neural Networks: Towards Generalization Across Time Steps and Deployment Friendliness
Kangrui Du, Yuhang Wu, Shikuang Deng +1
Spiking Neural Networks (SNNs), models inspired by neural mechanisms in the brain, allow for energy-efficient implementation on neuromorphic hardware. However, SNNs trained with cu…