6 papers
Mapping Overlaps in Benchmarks through Perplexity in the Wild
Siyang Wu, Honglin Bao, Sida Li +2
We introduce benchmark signatures to characterize the capacity demands of LLM benchmarks and their overlaps. Signatures are sets of salient tokens from in-the-wild corpora whose mo…
Language Models Should be Used to Surface the Unwritten Code of Science and Society
Honglin Bao, Siyang Wu, Jiwoong Choi +2
This paper calls on the research community not only to investigate how human biases are inherited by large language models (LLMs) but also to explore how these biases in LLMs can b…
Missing vs. Unused Knowledge Hypothesis for Language Model Bottlenecks in Patent Understanding
Siyang Wu, Honglin Bao, Nadav Kunievsky +1
While large language models (LLMs) excel at factual recall, the real challenge lies in knowledge application. A gap persists between their ability to answer complex questions and t…
Persistence Paradox in Dynamic Science
Honglin Bao, Kai Li
Persistence is often regarded as a virtue in science. In this paper, however, we challenge this conventional view by highlighting its contextual nature, particularly how persistenc…
Measuring Vogue in American Sociology (2011-2020)
Alex Xiaoqin Yan, Honglin Bao, Tom R. Leppard +1
This study investigates the social dynamics of knowledge production in American sociology. Departing from traditional approaches focused on citations, co-authorship, and faculty hi…
From Division to Unity: A Large-Scale Study on the Emergence of Computational Social Science, 1990-2021
Honglin Bao, Jiawei Zhang, Mingxuan Cao +1
We present a comprehensive study on the emergence of Computational Social Science (CSS) - an interdisciplinary field leveraging computational methods to address social science ques…