5 papers
Efficient Attention Mechanisms for Large Language Models: A Survey
Yutao Sun, Zhenyu Li, Yike Zhang +4
Transformer-based architectures have become the prevailing backbone of large language models. However, the quadratic time and memory complexity of self-attention remains a fundamen…
Negative Matters: Multi-Granularity Hard-Negative Synthesis and Anchor-Token-Aware Pooling for Enhanced Text Embeddings
Tengyu Pan, Zhichao Duan, Zhenyu Li +4
Text embedding models are essential for various natural language processing tasks, enabling the effective encoding of semantic information into dense vector representations. These…
Maximum Score Routing For Mixture-of-Experts
Bowen Dong, Yilong Fan, Yutao Sun +4
Routing networks in sparsely activated mixture-of-experts (MoE) dynamically allocate input tokens to top-k experts through differentiable sparse transformations, enabling scalable…
COMM:Concentrated Margin Maximization for Robust Document-Level Relation Extraction
Zhichao Duan, Tengyu Pan, Zhenyu Li +2
Document-level relation extraction (DocRE) is the process of identifying and extracting relations between entities that span multiple sentences within a document. Due to its realis…
TROI: Cross-Subject Pretraining with Sparse Voxel Selection for Enhanced fMRI Visual Decoding
Ziyu Wang, Tengyu Pan, Zhenyu Li +3
fMRI (functional Magnetic Resonance Imaging) visual decoding involves decoding the original image from brain signals elicited by visual stimuli. This often relies on manually label…