collaborators

5 papers

astro-ph.IM2026

AstroMLab 4: Benchmark-Topping Performance in Astronomy Q&A with a 70B-Parameter Domain-Specialized Reasoning Model

Tijmen de Haan, Yuan-Sen Ting, Tirthankar Ghosal +7

General-purpose large language models (LLMs), despite their broad capabilities, often struggle with specialized domain knowledge. This gap hinders their deployment as reliable rese…

cs.AI2026

A large-scale evaluation of commonsense knowledge in humans and large language models

Tuan Dung Nguyen, Duncan J. Watts, Mark E. Whiting

Commonsense knowledge, a major constituent of artificial intelligence (AI), is primarily evaluated in practice by human-prescribed ground-truth labels. An important, albeit implici…

astro-ph.IM2025

AstroMLab 5: Structured Summaries and Concept Extraction for 400,000 Astrophysics Papers

Yuan-Sen Ting, Alberto Accomazzi, Tirthankar Ghosal +4

We present a dataset of 408,590 astrophysics papers from arXiv (astro-ph), spanning 1992 through July 2025. Each paper has been processed through a multi-stage pipeline to produce:…

cs.CL2025

MoVa: Towards Generalizable Classification of Human Morals and Values

Ziyu Chen, Junfei Sun, Chenxi Li +6

Identifying human morals and values embedded in language is essential to empirical studies of communication. However, researchers often face substantial difficulty navigating the d…

astro-ph.IM2025

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model

Tijmen de Haan, Yuan-Sen Ting, Tirthankar Ghosal +6

AstroSage-Llama-3.1-8B is a domain-specialized natural-language AI assistant tailored for research in astronomy, astrophysics, cosmology, and astronomical instrumentation. Trained…