Publications (35)
AI of Brain and Cognitive Sciences: From the Perspective of First Principles
Luyao Chen, Zhiqiang Chen, Longsheng Jiang +13
Nowadays, we have witnessed the great success of AI in various applications, including image classification, game playing, protein structure analysis, language translation, and con…
Pre-training of Graph Neural Network for Modeling Effects of Mutations on Protein-Protein Binding Affinity
Xianggen Liu, Yunan Luo, Sen Song +1
Modeling the effects of mutations on the binding affinity plays a crucial role in protein engineering and drug design. In this study, we develop a novel deep learning based framewo…
Estimation of the volume of the left ventricle from MRI images using deep neural networks
Fangzhou Liao, Xi Chen, Xiaolin Hu +1
Segmenting human left ventricle (LV) in magnetic resonance imaging (MRI) images and calculating its volume are important for diagnosing cardiac diseases. In 2016, Kaggle organized…
Brain-Like Replay Naturally Emerges in Reinforcement Learning Agents
Jiyi Wang, Likai Tang, Huimiao Chen +2
Replay is a powerful strategy to promote learning in artificial intelligence and the brain. However, the conditions to generate it and its functional advantages have not been fully…
Contrastive Learning of Shared Spatiotemporal EEG Representations Across Individuals for Naturalistic Neuroscience
Xinke Shen, Lingyi Tao, Xuyang Chen +3
Neural representations induced by naturalistic stimuli offer insights into how humans respond to stimuli in daily life. Understanding neural mechanisms underlying naturalistic stim…
CaMKII activation supports reward-based neural network optimization through Hamiltonian sampling
Zhaofei Yu, David Kappel, Robert Legenstein +3
Synaptic plasticity is implemented and controlled through over thousand different types of molecules in the postsynaptic density and presynaptic boutons that assume a staggering ar…
Zooming Network
Yukun Yan, Daqi Zheng, Zhengdong Lu +1
Structural information is important in natural language understanding. Although some current neural net-based models have a limited ability to take local syntactic information, the…
Hierarchical Reasoning Model
Guan Wang, Jin Li, Yuhao Sun +6
Reasoning, the process of devising and executing complex goal-oriented action sequences, remains a critical challenge in AI. Current large language models (LLMs) primarily employ C…
EEG-JEPA: Structured Latent Prediction for EEG Foundation Models
Jinhao Li, Zhiyuan Ma, Xueqiao Han +8
Electroencephalography (EEG) foundation models aim to learn reusable representations from large-scale unlabeled recordings. A common pretraining strategy is masked waveform reconst…
Dynamic-Attention-based EEG State Transition Modeling for Emotion Recognition
Xinke Shen, Runmin Gan, Kaixuan Wang +5
Electroencephalogram (EEG)-based emotion decoding can objectively quantify people's emotional state and has broad application prospects in human-computer interaction and early dete…
Capturing Aperiodic Temporal Dynamics of EEG Signals through Stochastic Fluctuation Modeling
Yuhao Sun, Zhiyuan Ma, Xinke Shen +3
Electrophysiological brain signals, such as electroencephalography (EEG), exhibit both periodic and aperiodic components, with the latter often modeled as 1/f noise and considered…
Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models
Yixuan Liu, Zhiyuan Ma, Likai Tang +5
How large language models (LLMs) align with the neural representation and computation of human language is a central question in cognitive science. Using representational geometry…
Event Identification as a Decision Process with Non-linear Representation of Text
Yukun Yan, Daqi Zheng, Zhengdong Lu +1
We propose scale-free Identifier Network(sfIN), a novel model for event identification in documents. In general, sfIN first encodes a document into multi-scale memory stacks, then…
Signal-Adaptive Trust Regions for Gradient-Free Optimization of Recurrent Spiking Neural Networks
Jinhao Li, Yuhao Sun, Zhiyuan Ma +5
Recurrent spiking neural networks (RSNNs) are a promising substrate for energy-efficient control policies, but training them for high-dimensional, long-horizon reinforcement learni…
Evolving Connectivity for Recurrent Spiking Neural Networks
Guan Wang, Yuhao Sun, Sijie Cheng +1
Recurrent spiking neural networks (RSNNs) hold great potential for advancing artificial general intelligence, as they draw inspiration from the biological nervous system and show p…
OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
Guan Wang, Sijie Cheng, Xianyuan Zhan +3
Nowadays, open-source large language models like LLaMA have emerged. Recent developments have incorporated supervised fine-tuning (SFT) and reinforcement learning fine-tuning (RLFT…
Simulated annealing for optimization of graphs and sequences
Xianggen Liu, Pengyong Li, Fandong Meng +5
Optimization of discrete structures aims at generating a new structure with the better property given an existing one, which is a fundamental problem in machine learning. Different…
Learn molecular representations from large-scale unlabeled molecules for drug discovery
Pengyong Li, Jun Wang, Yixuan Qiao +6
How to produce expressive molecular representations is a fundamental challenge in AI-driven drug discovery. Graph neural network (GNN) has emerged as a powerful technique for model…
LI-DSN: A Layer-wise Interactive Dual-Stream Network for EEG Decoding
Chenghao Yue, Zhiyuan Ma, Zhongye Xia +4
Electroencephalography (EEG) provides a non-invasive window into brain activity, offering high temporal resolution crucial for understanding and interacting with neural processes t…
Local Hypergraph-based Nested Named Entity Recognition as Query-based Sequence Labeling
Yukun Yan, Sen Song
There has been a growing academic interest in the recognition of nested named entities in many domains. We tackle the task with a novel local hypergraph-based method: We first prop…
CRCC: Contrast-Based Robust Cross-Subject and Cross-Site Representation Learning for EEG
Xiaobin Wong, Zhonghua Zhao, Haoran Guo +5
EEG-based neural decoding models often fail to generalize across acquisition sites due to structured, site-dependent biases implicitly exploited during training. We reformulate cro…
Unsupervised Paraphrasing by Simulated Annealing
Xianggen Liu, Lili Mou, Fandong Meng +3
Unsupervised paraphrase generation is a promising and important research topic in natural language processing. We propose UPSA, a novel approach that accomplishes Unsupervised Para…
Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design
Sen Song
All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al. (2021) proved causes representational rank t…
UniMem: Towards a Unified View of Long-Context Large Language Models
Junjie Fang, Likai Tang, Hongzhe Bi +12
Long-context processing is a critical ability that constrains the applicability of large language models (LLMs). Although there exist various methods devoted to enhancing the long-…
JUMPER: Learning When to Make Classification Decisions in Reading
Xianggen Liu, Lili Mou, Haotian Cui +2
In early years, text classification is typically accomplished by feature-based machine learning models; recently, deep neural networks, as a powerful learning machine, make it poss…
Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision
Zhiyuan Ma, Zeyuan Li, Zhiyi Lu +7
The paper introduces BridgeMIL, a two-stage method that first learns EEG instance representations without using inherited labels and then applies subject-level supervision via a mu…
DSAINet: An Efficient Dual-Scale Attentive Interaction Network for General EEG Decoding
Zhiyuan Ma, Zeyuan Li, Zihao Qiu +6
In real-world applications of noninvasive electroencephalography (EEG), specialized decoders often show limited generalizability across diverse tasks under subject-independent sett…
KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG
Yongjian Li, HaoCheng Chu, Yukun Yan +7
Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to access broader knowledge sources, yet factual inconsistencies persist due to noise in retrieved documen…
Improve Decoding Factuality by Token-wise Cross Layer Entropy of Large Language Models
Jialiang Wu, Yi Shen, Sijia Liu +4
Despite their impressive capacities, Large language models (LLMs) often struggle with the hallucination issue of generating inaccurate or fabricated content even when they possess…
Evaluate the Malignancy of Pulmonary Nodules Using the 3D Deep Leaky Noisy-or Network
Fangzhou Liao, Ming Liang, Zhe Li +2
Automatic diagnosing lung cancer from Computed Tomography (CT) scans involves two steps: detect all suspicious lesions (pulmonary nodules) and evaluate the whole-lung/pulmonary mal…
Contrastive Learning of Subject-Invariant EEG Representations for Cross-Subject Emotion Recognition
Xinke Shen, Xianggen Liu, Xin Hu +2
EEG signals have been reported to be informative and reliable for emotion recognition in recent years. However, the inter-subject variability of emotion-related EEG signals still p…
Attentional Neural Network: Feature Selection Using Cognitive Feedback
Qian Wang, Jiaxing Zhang, Sen Song +1
Attentional Neural Network is a new framework that integrates top-down cognitive bias and bottom-up feature extraction in one coherent architecture. The top-down influence is espec…
Dual-Enhancement Product Bundling: Bridging Interactive Graph and Large Language Model
Zhe Huang, Peng Wang, Yan Zheng +2
Product bundling boosts e-commerce revenue by recommending complementary item combinations. However, existing methods face two critical challenges: (1) collaborative filtering appr…
Brain-inspired global-local learning incorporated with neuromorphic computing
Yujie Wu, Rong Zhao, Jun Zhu +11
Two main routes of learning methods exist at present including error-driven global learning and neuroscience-oriented local learning. Integrating them into one network may provide…
Why Attend to Everything? Focus is the Key
Hengshuai Yao, Xing Chen, Ahmed Murtadha +8
Standard attention scales quadratically with sequence length. Efficient attention methods reduce this O(n^2) cost, but when retrofitted into pretrained models, they often degrade p…