23 papers
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents
Yujun Zhou, Kehan Guo, Haomin Zhuang +8
Interactive LLM agents are becoming part of daily work, but they do not reliably become easier to work with over time: a correction remembered in one session may still be violated…
Alignment Risks from Capability-Seeking RL Training
Yujun Zhou, Yue Huang, Han Bao +8
While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerabl…
Genotype-Conditioned Molecular Generation via Evidence-Grounded Multi-Objective Latent Perturbation in Diffusion Models
Brenda Nogueira, Gisela A. Gonzalez-Montiel, Nitesh V. Chawla +1
Developing effective anticancer therapeutics remains challenging due to tumor heterogeneity and the absence of well-defined molecular targets across cancer subtypes. Generative mod…
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
Ngoc Phan Phuoc Loc, Toan Huynh La Viet, Thanh Tran Khanh +8
The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automated peer reviewers. However, h…
SPECTRA: Spectral Domain-Aware Graph Generation for Imbalanced Molecular Property Regression
Brenda Nogueira, Gisela A. Gonzalez-Montiel, Meng Jiang +2
Molecular property regression struggles with cases in chemically relevant target ranges that are underrepresented in datasets. Standard average error minimization approaches underp…
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
Kaiwen Shi, Weixiang Sun, Zheyuan Zhang +3
Scientific research relies on citation integrity, yet large language models (LLMs) have introduced a critical risk: fabricated references that appear plausible but correspond to no…