activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2024

HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation

Yuhan Chen, Ang Lv, Jian Luan +2

Many positional encodings (PEs) are designed to exhibit long-term decay, based on an entrenched and long-standing inductive opinion: tokens farther away from the current position c…

cs.CL2024

An Analysis and Mitigation of the Reversal Curse

Ang Lv, Kaiyi Zhang, Shufang Xie +4

Recent research observed a noteworthy phenomenon in large language models (LLMs), referred to as the ``reversal curse.'' The reversal curse is that when dealing with two entities,…

cs.CL2024

Mixture of In-Context Experts Enhance LLMs' Long Context Awareness

Hongzhan Lin, Ang Lv, Yuhan Chen +4

Many studies have revealed that large language models (LLMs) exhibit uneven awareness of different contextual positions. Their limited context awareness can lead to overlooking cri…

cs.CL2024

YuLan: An Open-source Large Language Model

Yutao Zhu, Kun Zhou, Kelong Mao +35

Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many o…

cs.CL2024

Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use

Yuhan Chen, Ang Lv, Ting-En Lin +5

In this paper, we demonstrate that an inherent waveform pattern in the attention allocation of large language models (LLMs) significantly affects their performance in tasks demandi…

cs.CL2024

Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models

Ang Lv, Yuhan Chen, Kaiyi Zhang +5

In this paper, we delve into several mechanisms employed by Transformer-based language models (LLMs) for factual recall tasks. We outline a pipeline consisting of three major steps…