27 papers
Why AI Detection Fails for Academic Integrity
Jonathan A. Karr, Grigorii Khvatskii, Ting Hua +1
Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled…
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Peiyu Li, Xiuxiu Tang, Si Chen +4
The paper proposes ATLAS, an adaptive testing framework using Item Response Theory to evaluate large language models more efficiently by selecting informative items, reducing requi…
Policy4OOD: A Knowledge-Guided World Model for Policy Intervention Simulation against the Opioid Overdose Crisis
Yijun Ma, Zehong Wang, Weixiang Sun +4
The opioid epidemic remains one of the most severe public health crises in the United States, yet evaluating policy interventions before implementation is difficult: multiple polic…
Can Decision Trees Teach Large Language Models? Distilling Verbalized Knowledge for Molecular Property Prediction
Khiem Le, Sreejata Dey, Marcos MartÃnez Galindo +4
Molecular Property Prediction (MPP) is a fundamental problem in drug discovery that has recently attracted growing attention. Large Language Models (LLMs), known for their impressi…
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
Khiem Le, Phuc Nguyen, Youssef Mroueh +4
Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models, but it suffers from two critic…
Do LLMs have core beliefs?
Anna Sokol, Marianna B. Ganapini, Nitesh V. Chawla
The rise of Large Language Models (LLMs) has sparked debate about whether these systems exhibit human-level cognition. In this debate, little attention has been paid to a structura…