2 papers
cs.AI2026
ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
Matthew ffrench-Constant, Daniel Yang, Xinmeng Huang +1
Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone is insufficient: models must…
cs.AI2024
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
Xinmeng Huang, Shuo Li, Edgar Dobriban +3
The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safe…