collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts

Xiaoqiang Wang, Suyuchen Wang, Yun Zhu +1

Chain-of-thought (CoT) reasoning enables large language models (LLMs) to move beyond fast System-1 responses and engage in deliberative System-2 reasoning. However, this comes at t…

cs.CL2025

RMem: Bridging Memory Retention and Retrieval via Reversible Compression

Xiaoqiang Wang, Suyuchen Wang, Yun Zhu +1

Memory plays a key role in enhancing LLMs' performance when deployed to real-world applications. Existing solutions face trade-offs: explicit memory designs based on external stora…

cs.CL2024

FACT: Examining the Effectiveness of Iterative Context Rewriting for Multi-fact Retrieval

Jinlin Wang, Suyuchen Wang, Ziwen Xia +4

Large Language Models (LLMs) are proficient at retrieving single facts from extended contexts, yet they struggle with tasks requiring the simultaneous retrieval of multiple facts,…

cs.CL2024

LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models

Zhiyuan Hu, Yuliang Liu, Jinman Zhao +8

Large language models (LLMs) face significant challenges in handling long-context tasks because of their limited effective context window size during pretraining, which restricts t…

cs.CL2024

Resonance RoPE: Improving Context Length Generalization of Large Language Models

Suyuchen Wang, Ivan Kobyzev, Peng Lu +2

This paper addresses the challenge of train-short-test-long (TSTL) scenarios in Large Language Models (LLMs) equipped with Rotary Position Embedding (RoPE), where models pre-traine…