4 papers
Label-Free Reinforcement Learning via Cross-Model Entropy
Matt Gorbett, Hossein Shirazi
Post-training large language models with reinforcement learning is bottlenecked by the reward signal. Existing approaches require either ground-truth verifiable rewards, restrictin…
Cross-Model Disagreement as a Label-Free Correctness Signal
Matt Gorbett, Suman Jana
Detecting when a language model is wrong without ground truth labels is a fundamental challenge for safe deployment. Existing approaches rely on a model's own uncertainty -- such a…
Characterizing Linear Alignment Across Language Models
Matt Gorbett, Suman Jana
Language models increasingly appear to learn similar representations, despite differences in training objectives, architectures, and data modalities. This emerging compatibility be…
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
Matt Gorbett, Hossein Shirazi, Indrakshi Ray
Binary Neural Networks (BNNs) enable efficient deep learning by saving on storage and computational costs. However, as the size of neural networks continues to grow, meeting comput…