5 papers
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
Shengrui Li, Fei Zhao, Kaiyan Zhao +6
Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence with proficiency on hard tasks such a…
Balancing Understanding and Generation in Discrete Diffusion Models
Yue Liu, Yuzhong Zhao, Zheyong Xie +5
In discrete generative modeling, two dominant paradigms demonstrate divergent capabilities: Masked Diffusion Language Models (MDLM) excel at semantic understanding and zero-shot ge…
Benchmarking Machine Translation on Chinese Social Media Texts
Kaiyan Zhao, Zheyong Xie, Zhongtao Miao +3
The prevalence of rapidly evolving slang, neologisms, and highly stylized expressions in informal user-generated text, particularly on Chinese social media, poses significant chall…
EComStage: Stage-wise and Orientation-specific Benchmarking for Large Language Models in E-commerce
Kaiyan Zhao, Zijie Meng, Zheyong Xie +4
Large Language Model (LLM)-based agents are increasingly deployed in e-commerce applications to assist customer services in tasks such as product inquiries, recommendations, and or…
RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
Fei Zhao, Chonggang Lu, Haofu Qian +9
As a key medium for human interaction and information exchange, social networking services (SNS) pose unique challenges for large language models (LLMs): heterogeneous workloads, f…