6 papers
FinS-Pilot: A Benchmark for Online Financial RAG System
Feng Wang, Yiding Sun, Jiaxin Mao +2
Large language models (LLMs) have demonstrated remarkable capabilities across various professional domains, with their performance typically evaluated through standardized benchmar…
BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baseline
Guosheng Dong, Da Pan, Yiding Sun +17
The general capabilities of Large Language Models (LLM) highly rely on the composition and selection on extensive pretraining datasets, treated as commercial secrets by several ins…
Aligning Explanations for Recommendation with Rating and Feature via Maximizing Mutual Information
Yurou Zhao, Yiding Sun, Ruidong Han +6
Providing natural language-based explanations to justify recommendations helps to improve users' satisfaction and gain users' trust. However, as current explanation generation meth…
Towards Effective and Efficient Continual Pre-training of Large Language Models
Jie Chen, Zhipeng Chen, Jiapeng Wang +16
Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents…
YuLan: An Open-source Large Language Model
Yutao Zhu, Kun Zhou, Kelong Mao +35
Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many o…
An Integrated Data Processing Framework for Pretraining Foundation Models
Yiding Sun, Feng Wang, Yutao Zhu +2
The ability of the foundation models heavily relies on large-scale, diverse, and high-quality pretraining data. In order to improve data quality, researchers and practitioners ofte…