3 papers
cs.AI2026
Cross-Model Disagreement as a Label-Free Correctness Signal
Matt Gorbett, Suman Jana
Detecting when a language model is wrong without ground truth labels is a fundamental challenge for safe deployment. Existing approaches rely on a model's own uncertainty -- such a…
cs.LG2026
Label-Free Reinforcement Learning via Cross-Model Entropy
Matt Gorbett, Hossein Shirazi
Post-training large language models with reinforcement learning is bottlenecked by the reward signal. Existing approaches require either ground-truth verifiable rewards, restrictin…
cs.AI2026
Characterizing Linear Alignment Across Language Models
Matt Gorbett, Suman Jana
Language models increasingly appear to learn similar representations, despite differences in training objectives, architectures, and data modalities. This emerging compatibility be…