activity
20242026
most citedAstroMLab 2: AstroLLaMA-2-70B Model and Benchmarking Specialised LLMs for Astronomy

2 citations · 2 across the 3 of their papers we have counts for

collaborators

5 papers

cs.SE2026

Investigating Tool-Memory Conflicts in Tool-Augmented LLMs

Jiali Cheng, Rui Pan, Hadi Amiri

Tool-augmented large language models (LLMs) have powered many applications. However, they are likely to suffer from knowledge conflict. In this paper, we propose a new type of know…

astro-ph.IM2025

AstroMLab 5: Structured Summaries and Concept Extraction for 400,000 Astrophysics Papers

Yuan-Sen Ting, Alberto Accomazzi, Tirthankar Ghosal +4

We present a dataset of 408,590 astrophysics papers from arXiv (astro-ph), spanning 1992 through July 2025. Each paper has been processed through a multi-stage pipeline to produce:…

cs.CL2025

Subjective Perspectives within Learned Representations Predict High-Impact Innovation

Likun Cao, Rui Pan, James Evans

Existing studies of innovation emphasize the power of social structures to shape innovation capacity. Emerging machine learning approaches, however, enable us to model innovators'…

astro-ph.IM2024

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model

Tijmen de Haan, Yuan-Sen Ting, Tirthankar Ghosal +6

AstroSage-Llama-3.1-8B is a domain-specialized natural-language AI assistant tailored for research in astronomy, astrophysics, cosmology, and astronomical instrumentation. Trained…

astro-ph.IM20242 cited

AstroMLab 2: AstroLLaMA-2-70B Model and Benchmarking Specialised LLMs for Astronomy

Rui Pan, Tuan Dung Nguyen, Hardik Arora +3

Continual pretraining of large language models on domain-specific data has been proposed to enhance performance on downstream tasks. In astronomy, the previous absence of astronomy…