2 papers
cs.CL2025
Diverse Preference Learning for Capabilities and Alignment
Stewart Slocum, Asher Parker-Sartori, Dylan Hadfield-Menell
The ability of LLMs to represent diverse perspectives is critical as they increasingly impact society. However, recent studies reveal that alignment algorithms such as RLHF and DPO…
cs.LG2025
On the creation of narrow AI: hierarchy and nonlocality of neural network skills
Eric J. Michaud, Asher Parker-Sartori, Max Tegmark
We study the problem of creating strong, yet narrow, AI systems. While recent AI progress has been driven by the training of large general-purpose foundation models, the creation o…