4 papers
Unlocking the Pre-Trained Model as a Dual-Alignment Calibrator for Post-Trained LLMs
Beier Luo, Cheng Wang, Hongxin Wei +2
Post-training improves large language models (LLMs) but often worsens confidence calibration, leading to systematic overconfidence. Recent unsupervised post-hoc methods for post-tr…
A Definition of AGI
Dan Hendrycks, Dawn Song, Christian Szegedy +30
The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quant…
How Well Can Preference Optimization Generalize Under Noisy Feedback?
Shawn Im, Sharon Li
As large language models (LLMs) advance their capabilities, aligning these models with human preferences has become crucial. Preference optimization, which trains models to disting…
Harnessing Feature Resonance under Arbitrary Target Alignment for Out-of-Distribution Node Detection
Shenzhi Yang, Junbo Zhao, Sharon Li +4
Detecting out-of-distribution (OOD) nodes in the graph-based machine-learning field is challenging, particularly when in-distribution (ID) node multi-category labels are unavailabl…