Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
Chejian Xu, Wei Ping, Peng Xu +5
Long-context capabilities are essential for a wide range of applications, including document and video understanding, in-context learning, and inference-time scaling, all of which…
cs.CL2023
InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
Boxin Wang, Wei Ping, Lawrence McAfee +4
Pretraining auto-regressive large language models~(LLMs) with retrieval demonstrates better perplexity and factual accuracy by leveraging external databases. However, the size of e…