8 papers
Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design
Sen Song
All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al. (2021) proved causes representational rank t…
EEG-JEPA: Structured Latent Prediction for EEG Foundation Models
Jinhao Li, Zhiyuan Ma, Xueqiao Han +8
Electroencephalography (EEG) foundation models aim to learn reusable representations from large-scale unlabeled recordings. A common pretraining strategy is masked waveform reconst…
Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision
Zhiyuan Ma, Zeyuan Li, Zhiyi Lu +7
The paper introduces BridgeMIL, a two-stage method that first learns EEG instance representations without using inherited labels and then applies subject-level supervision via a mu…
Why Attend to Everything? Focus is the Key
Hengshuai Yao, Xing Chen, Ahmed Murtadha +8
Standard attention scales quadratically with sequence length. Efficient attention methods reduce this O(n^2) cost, but when retrofitted into pretrained models, they often degrade p…
Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models
Yixuan Liu, Zhiyuan Ma, Likai Tang +5
How large language models (LLMs) align with the neural representation and computation of human language is a central question in cognitive science. Using representational geometry…
Signal-Adaptive Trust Regions for Gradient-Free Optimization of Recurrent Spiking Neural Networks
Jinhao Li, Yuhao Sun, Zhiyuan Ma +5
Recurrent spiking neural networks (RSNNs) are a promising substrate for energy-efficient control policies, but training them for high-dimensional, long-horizon reinforcement learni…