1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.AI2026
Laguna M.1/XS.2 Technical Report
Julien Abadji, Marah Abdin, Connor Adams +93
We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has B total parameters (B activated per tok…
cs.LG2025
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
Yiran Shen, Yu Xia, Jonathan Chang +1
Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a single objective. We seek to answer wha…
cs.LG2024★ 1 cited
Critique-out-Loud Reward Models
Zachary Ankner, Mansheej Paul, Brandon Cui +2
Traditionally, reward models used for reinforcement learning from human feedback (RLHF) are trained to directly predict preference scores without leveraging the generation capabili…