collaborators

6 papers

cs.CL2025

FinS-Pilot: A Benchmark for Online Financial RAG System

Feng Wang, Yiding Sun, Jiaxin Mao +2

Large language models (LLMs) have demonstrated remarkable capabilities across various professional domains, with their performance typically evaluated through standardized benchmar…

cs.CL2024

BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baseline

Guosheng Dong, Da Pan, Yiding Sun +17

The general capabilities of Large Language Models (LLM) highly rely on the composition and selection on extensive pretraining datasets, treated as commercial secrets by several ins…

cs.IR2024

Aligning Explanations for Recommendation with Rating and Feature via Maximizing Mutual Information

Yurou Zhao, Yiding Sun, Ruidong Han +6

Providing natural language-based explanations to justify recommendations helps to improve users' satisfaction and gain users' trust. However, as current explanation generation meth…

cs.CL2024

Towards Effective and Efficient Continual Pre-training of Large Language Models

Jie Chen, Zhipeng Chen, Jiapeng Wang +16

Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents…

cs.CL2024

YuLan: An Open-source Large Language Model

Yutao Zhu, Kun Zhou, Kelong Mao +35

Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many o…

cs.LG2024

An Integrated Data Processing Framework for Pretraining Foundation Models

Yiding Sun, Feng Wang, Yutao Zhu +2

The ability of the foundation models heavily relies on large-scale, diverse, and high-quality pretraining data. In order to improve data quality, researchers and practitioners ofte…