collaborators

27 papers

cs.LG2026

Why AI Detection Fails for Academic Integrity

Jonathan A. Karr, Grigorii Khvatskii, Ting Hua +1

Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled…

cs.CL2026

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Peiyu Li, Xiuxiu Tang, Si Chen +4

The paper proposes ATLAS, an adaptive testing framework using Item Response Theory to evaluate large language models more efficiently by selecting informative items, reducing requi…

cs.LG2026

Policy4OOD: A Knowledge-Guided World Model for Policy Intervention Simulation against the Opioid Overdose Crisis

Yijun Ma, Zehong Wang, Weixiang Sun +4

The opioid epidemic remains one of the most severe public health crises in the United States, yet evaluating policy interventions before implementation is difficult: multiple polic…

cs.LG2026

Can Decision Trees Teach Large Language Models? Distilling Verbalized Knowledge for Molecular Property Prediction

Khiem Le, Sreejata Dey, Marcos Martínez Galindo +4

Molecular Property Prediction (MPP) is a fundamental problem in drug discovery that has recently attracted growing attention. Large Language Models (LLMs), known for their impressi…

cs.LG2026

Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models

Khiem Le, Phuc Nguyen, Youssef Mroueh +4

Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models, but it suffers from two critic…

cs.LG2026

Do LLMs have core beliefs?

Anna Sokol, Marianna B. Ganapini, Nitesh V. Chawla

The rise of Large Language Models (LLMs) has sparked debate about whether these systems exhibit human-level cognition. In this debate, little attention has been paid to a structura…