collaborators

5 papers

cs.CL2026

Efficient Attention Mechanisms for Large Language Models: A Survey

Yutao Sun, Zhenyu Li, Yike Zhang +4

Transformer-based architectures have become the prevailing backbone of large language models. However, the quadratic time and memory complexity of self-attention remains a fundamen…

cs.CL2025

Negative Matters: Multi-Granularity Hard-Negative Synthesis and Anchor-Token-Aware Pooling for Enhanced Text Embeddings

Tengyu Pan, Zhichao Duan, Zhenyu Li +4

Text embedding models are essential for various natural language processing tasks, enabling the effective encoding of semantic information into dense vector representations. These…

cs.LG2025

Maximum Score Routing For Mixture-of-Experts

Bowen Dong, Yilong Fan, Yutao Sun +4

Routing networks in sparsely activated mixture-of-experts (MoE) dynamically allocate input tokens to top-k experts through differentiable sparse transformations, enabling scalable…

cs.CL2025

COMM:Concentrated Margin Maximization for Robust Document-Level Relation Extraction

Zhichao Duan, Tengyu Pan, Zhenyu Li +2

Document-level relation extraction (DocRE) is the process of identifying and extracting relations between entities that span multiple sentences within a document. Due to its realis…

cs.CV2025

TROI: Cross-Subject Pretraining with Sparse Voxel Selection for Enhanced fMRI Visual Decoding

Ziyu Wang, Tengyu Pan, Zhenyu Li +3

fMRI (functional Magnetic Resonance Imaging) visual decoding involves decoding the original image from brain signals elicited by visual stimuli. This often relies on manually label…