3 papers
cs.AI2026
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs
Darsh Kachroo, Arjun Prasaath Anbazhagan, Adriana Caraeni +2
Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex reasoning tasks. DPO optimiz…
cs.LG2026
Predicting Groundwater Arsenic Concentrations Using Graph Neural Networks
William Xing, Stephanie Yang, Aarush Bandemegal +4
Arsenic contamination in groundwater presents a longstanding public health crisis in the United States, especially for households depending on private wells. Accurate and spatially…
math.AT2025
Exploring the Stratified Space Structure of an RL Game with the Volume Growth Transform
Justin Curry, Brennan Lagasse, Ngoc B. Lam +3
In this work, we explore the structure of the embedding space of a transformer model trained for playing a particular reinforcement learning (RL) game. Specifically, we investigate…