7 papers · 1 filter
BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language
Qizhi Pei, Zhimeng Zhou, Yi Duan +9
We present BioMatrix, the first multimodal foundation model that natively integrates sequences, structures, and natural language for both molecules and proteins within a single dec…
Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
Songhao Wu, Zhongxin Chen, Yuxuan Liu +3
Large language models exhibit impressive zero-shot capabilities across a wide range of downstream tasks. However, they struggle to function as off-the-shelf embedding models, leadi…
HierBias: Context-Conditioned Hierarchical Media Bias Detection with Multi-Task Type Classification
Kaining Li, Ruichen Yan, Yuxin Dong
Media bias detection is a critical task for ensuring fair and balanced information dissemination, yet existing sentence-level approaches classify each sentence independently, ignor…
Mixture of In-Context Experts Enhance LLMs' Long Context Awareness
Hongzhan Lin, Ang Lv, Yuhan Chen +4
Many studies have revealed that large language models (LLMs) exhibit uneven awareness of different contextual positions. Their limited context awareness can lead to overlooking cri…
Towards Effective and Efficient Continual Pre-training of Large Language Models
Jie Chen, Zhipeng Chen, Jiapeng Wang +16
Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents…
YuLan: An Open-source Large Language Model
Yutao Zhu, Kun Zhou, Kelong Mao +35
Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many o…