7 papers · 1 filter
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
Ziyang Chen, Xing Wu, Junlong Jia +4
The rapid expansion of context length in large language models (LLMs) has outpaced existing evaluation benchmarks. Current long-context benchmarks often trade off scalability and r…
EntropyLong: Effective Long-Context Training via Predictive Uncertainty
Junlong Jia, Ziyang Chen, Xing Wu +5
Training long-context language models to capture long-range dependencies requires specialized data construction. Current approaches, such as generic text concatenation or heuristic…
LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
Junlong Jia, Xing Wu, Chaochen Gao +8
High-quality long-context data is essential for training large language models (LLMs) capable of processing extensive documents, yet existing synthesis approaches using relevance-b…
LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
Chaochen Gao, Xing Wu, Zijia Lin +2
High-quality long-context instruction data is essential for aligning long-context large language models (LLMs). Despite the public release of models like Qwen and Llama, their long…
TransAug: Translate as Augmentation for Sentence Embeddings
Jue Wang
While contrastive learning greatly advances the representation of sentence embeddings, it is still limited by the size of the existing sentence datasets. In this paper, we present…
NExtLong: Toward Effective Long-Context Training without Long Documents
Chaochen Gao, Xing Wu, Zijia Lin +2
Large language models (LLMs) with extended context windows have made significant strides yet remain a challenge due to the scarcity of long documents. Existing methods tend to synt…