Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Model Merging on Loss Landscape: A Geometry Perspective
Juanwu Lu, Anand Bhaskar, Brian Axelrod +2
Model merging offers a promising avenue for knowledge integration and parallel development without retraining. Yet, existing methods either ignore the geometry of the loss landscap…
cs.LG2026
Attention Head Entropy of LLMs Predicts Answer Correctness
Sophie Ostmeier, Brian Axelrod, Maya Varma +6
Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-ju…
cs.LG2024
Sample Amplification: Increasing Dataset Size even when Learning is Impossible
Brian Axelrod, Shivam Garg, Vatsal Sharan +1
Given data drawn from an unknown distribution, , to what extent is it possible to ``amplify'' this dataset and output an even larger set of samples that appear to have been draw…