activity
20242026
collaborators

6 papers

cs.AI2026

Mapping Overlaps in Benchmarks through Perplexity in the Wild

Siyang Wu, Honglin Bao, Sida Li +2

We introduce benchmark signatures to characterize the capacity demands of LLM benchmarks and their overlaps. Signatures are sets of salient tokens from in-the-wild corpora whose mo…

cs.CY2026

Language Models Should be Used to Surface the Unwritten Code of Science and Society

Honglin Bao, Siyang Wu, Jiwoong Choi +2

This paper calls on the research community not only to investigate how human biases are inherited by large language models (LLMs) but also to explore how these biases in LLMs can b…

cs.CL2026

Missing vs. Unused Knowledge Hypothesis for Language Model Bottlenecks in Patent Understanding

Siyang Wu, Honglin Bao, Nadav Kunievsky +1

While large language models (LLMs) excel at factual recall, the real challenge lies in knowledge application. A gap persists between their ability to answer complex questions and t…

cs.DL2025

Persistence Paradox in Dynamic Science

Honglin Bao, Kai Li

Persistence is often regarded as a virtue in science. In this paper, however, we challenge this conventional view by highlighting its contextual nature, particularly how persistenc…

cs.SI2025

Measuring Vogue in American Sociology (2011-2020)

Alex Xiaoqin Yan, Honglin Bao, Tom R. Leppard +1

This study investigates the social dynamics of knowledge production in American sociology. Departing from traditional approaches focused on citations, co-authorship, and faculty hi…

cs.CY2024

From Division to Unity: A Large-Scale Study on the Emergence of Computational Social Science, 1990-2021

Honglin Bao, Jiawei Zhang, Mingxuan Cao +1

We present a comprehensive study on the emergence of Computational Social Science (CSS) - an interdisciplinary field leveraging computational methods to address social science ques…