3 papers
cs.LG2026
Model Merging on Loss Landscape: A Geometry Perspective
Juanwu Lu, Anand Bhaskar, Brian Axelrod +2
Model merging offers a promising avenue for knowledge integration and parallel development without retraining. Yet, existing methods either ignore the geometry of the loss landscap…
cs.LG2026
Attention Head Entropy of LLMs Predicts Answer Correctness
Sophie Ostmeier, Brian Axelrod, Maya Varma +6
Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-ju…
cs.CV2025
LieRE: Lie Rotational Positional Encodings
Sophie Ostmeier, Brian Axelrod, Maya Varma +3
Transformer architectures rely on position encodings to model the spatial structure of input data. Rotary Position Encoding (RoPE) is a widely used method in language models that e…