8 papers
StatCite: A Large-scale Citation Network Dataset for Statistics and Data Science
Tianang Deng, Tianchen Gao, Rui Pan +1
In this paper, we introduce StatCite, a large-scale citation network dataset covering publications in statistics and data science from 1981 to 2025. The dataset contains 189,101 re…
An LLM-Powered Semantic Alignment Framework for Journal Recommendation
Yanglin Yan, Zicheng Xie, Tianchen Gao +2
Journal recommendation is an important task in scholarly information systems. Existing approaches typically rely on supervised learning models, manually engineered features, or his…
How Does LLM Help Regional CPI Forecast: An LLM-powered Deep Panel Modeling Framework
Tianchen Gao, Ao Sun, Yurou Wang +2
Understanding regional Consumer Price Index (CPI) dynamics is essential for timely and effective economic policymaking. However, traditional modeling procedures typically rely only…
LLM-powered Real-time Patent Citation Recommendation for Financial Technologies
Tianang Deng, Yu Deng, Tianchen Gao +2
Rapid financial innovation has been accompanied by a sharp increase in patenting activity, making timely and comprehensive prior-art discovery more difficult. This problem is espec…
Are Large Language Models able to Predict Highly Cited Papers? Evidence from Statistical Publications
Zhanshuo Ye, Yiming Hou, Rui Pan +2
Predicting highly-cited papers is a long-standing challenge due to the complex interactions of research content, scholarly communities, and temporal dynamics. Recent advances in la…
A Comparison of DeepSeek and Other LLMs
Tianchen Gao, Jiashun Jin, Zheng Tracy Ke +1
Recently, DeepSeek has been the focus of attention in and beyond the AI community. An interesting problem is how DeepSeek compares to other large language models (LLMs). There are…