activity
20242026
collaborators

10 papers

cs.CL2026

MiLe Loss: a New Entropy-Weighed Loss for Mitigating the Bias of Learning Difficulties in Large Language Models

Zhenpeng Su, Xing Wu, Xue Bai +5

Generative language models are usually pretrained on large text corpus via predicting the next token (i.e., sub-word/word/phrase) given the previous ones. Recent works have demonst…

cs.CL2025

EntropyLong: Effective Long-Context Training via Predictive Uncertainty

Junlong Jia, Ziyang Chen, Xing Wu +5

Training long-context language models to capture long-range dependencies requires specialized data construction. Current approaches, such as generic text concatenation or heuristic…

cs.AI2025

Libra: Large Chinese-based Safeguard for AI Content

Ziyang Chen, Huimu Yu, Xing Wu +2

Large language models (LLMs) excel in text understanding and generation but raise significant safety and ethical concerns in high-stakes applications. To mitigate these risks, we p…

cs.CL2025

LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions

Chaochen Gao, Xing Wu, Zijia Lin +2

High-quality long-context instruction data is essential for aligning long-context large language models (LLMs). Despite the public release of models like Qwen and Llama, their long…

cs.AI2025

CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoning

Huimu Yu, Xing Wu, Haotian Xu +2

Large language models (LLMs) have made significant progress in natural language understanding and generation, driven by scalable pretraining and advanced finetuning. However, enhan…

cs.CL2025

NExtLong: Toward Effective Long-Context Training without Long Documents

Chaochen Gao, Xing Wu, Zijia Lin +2

Large language models (LLMs) with extended context windows have made significant strides yet remain a challenge due to the scarcity of long documents. Existing methods tend to synt…