Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Cross-Model Disagreement as a Label-Free Correctness Signal
Matt Gorbett, Suman Jana
Detecting when a language model is wrong without ground truth labels is a fundamental challenge for safe deployment. Existing approaches rely on a model's own uncertainty -- such a…
cs.AI2026
Characterizing Linear Alignment Across Language Models
Matt Gorbett, Suman Jana
Language models increasingly appear to learn similar representations, despite differences in training objectives, architectures, and data modalities. This emerging compatibility be…