14 papers
Why AI Detection Fails for Academic Integrity
Jonathan A. Karr, Grigorii Khvatskii, Ting Hua +1
Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled…
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Peiyu Li, Xiuxiu Tang, Si Chen +4
The paper proposes ATLAS, an adaptive testing framework using Item Response Theory to evaluate large language models more efficiently by selecting informative items, reducing requi…
Can Decision Trees Teach Large Language Models? Distilling Verbalized Knowledge for Molecular Property Prediction
Khiem Le, Sreejata Dey, Marcos MartÃnez Galindo +4
Molecular Property Prediction (MPP) is a fundamental problem in drug discovery that has recently attracted growing attention. Large Language Models (LLMs), known for their impressi…
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
Khiem Le, Phuc Nguyen, Youssef Mroueh +4
Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models, but it suffers from two critic…
TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
Isabel Molnar, Peiyu Li, Si Chen +5
Higher education instructors often lack timely and pedagogically grounded support, as scalable instructional guidance remains limited and existing tools rely on generic chatbot adv…
AgentDrug: Utilizing Large Language Models in An Agentic Workflow for Zero-Shot Molecular Editing
Khiem Le, Ting Hua, Nitesh V. Chawla
Molecular editing-modifying a given molecule to improve desired properties-is a fundamental task in drug discovery. While LLMs hold the potential to solve this task using natural l…