9 papers
Mitigating LLM-based p-Hacking by Preregistering for the Next LLM
Maria Thomas, Kristina Gligoric, Nihar B. Shah
Large language models (LLMs) are increasingly used to generate, classify, and annotate data whose outputs feed downstream hypothesis tests. However, LLM-based research is easy to p…
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
Rounak Saha, Gurusha Juneja, Dayita Chaudhuri +3
A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction…
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
Shanda Li, Qiuhong Anna Wei, Jingwu Tang +5
Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to evaluate whether LLM agents can assist wi…
How Many Submissions May an Author Make? A Harmonic Quota for Submissions under Coauthorship
Nihar B. Shah
Research evaluation systems -- including journals, conferences, and funders -- are increasingly using author-level submission limits to manage growing submission loads. Most existi…
Smooth Partial Lotteries for Stable Randomized Selection
Alexander Goldberg, Giulia Fanti, Nihar B. Shah
Competitive selection processes, from scientific funding to admissions and hiring, use evaluations to score candidates, and eventually choose a subset of them based on those scores…
Learning What Evaluators Value: A Reliable Approach to Modeling Evaluator Preferences
Madeline Celi Kitch, Nihar B. Shah
In many applications, human and LLM evaluators use assessments of relevant criteria to create an overall evaluation for an item or individual. For example, in admissions, committee…