collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark

Ziyang Chen, Xing Wu, Junlong Jia +4

The rapid expansion of context length in large language models (LLMs) has outpaced existing evaluation benchmarks. Current long-context benchmarks often trade off scalability and r…

cs.CL2025

EntropyLong: Effective Long-Context Training via Predictive Uncertainty

Junlong Jia, Ziyang Chen, Xing Wu +5

Training long-context language models to capture long-range dependencies requires specialized data construction. Current approaches, such as generic text concatenation or heuristic…

cs.CL2025

LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs

Junlong Jia, Xing Wu, Chaochen Gao +8

High-quality long-context data is essential for training large language models (LLMs) capable of processing extensive documents, yet existing synthesis approaches using relevance-b…

cs.CL2025

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions

Chaochen Gao, Xing Wu, Zijia Lin +2

High-quality long-context instruction data is essential for aligning long-context large language models (LLMs). Despite the public release of models like Qwen and Llama, their long…

cs.CL2025

TransAug: Translate as Augmentation for Sentence Embeddings

Jue Wang

While contrastive learning greatly advances the representation of sentence embeddings, it is still limited by the size of the existing sentence datasets. In this paper, we present…

cs.CL2025

NExtLong: Toward Effective Long-Context Training without Long Documents

Chaochen Gao, Xing Wu, Zijia Lin +2

Large language models (LLMs) with extended context windows have made significant strides yet remain a challenge due to the scarcity of long documents. Existing methods tend to synt…