4 papers
Why Attend to Everything? Focus is the Key
Hengshuai Yao, Xing Chen, Ahmed Murtadha +8
Standard attention scales quadratically with sequence length. Efficient attention methods reduce this O(n^2) cost, but when retrofitted into pretrained models, they often degrade p…
An Evolutionary Algorithm with Probabilistic Annealing for Large-scale Sparse Multi-objective Optimization
Shuai Shao, Yuhao Sun, Xing Chen +3
Large-scale sparse multi-objective optimization problems (LSMOPs) are prevalent in real-world applications, where optimal solutions typically contain only a few nonzero variables,…
Signal-Adaptive Trust Regions for Gradient-Free Optimization of Recurrent Spiking Neural Networks
Jinhao Li, Yuhao Sun, Zhiyuan Ma +5
Recurrent spiking neural networks (RSNNs) are a promising substrate for energy-efficient control policies, but training them for high-dimensional, long-horizon reinforcement learni…
Hierarchical Reasoning Model
Guan Wang, Jin Li, Yuhao Sun +6
Reasoning, the process of devising and executing complex goal-oriented action sequences, remains a critical challenge in AI. Current large language models (LLMs) primarily employ C…