2 papers
cs.LG2026
Spectral Souping: A Unified Framework for Online Preference Alignment
Yinlam Chow, Guy Tennenholtz, Ted Yun +4
Reinforcement Learning from Human Feedback (RLHF) effectively aligns Large Language Models (LLMs) with aggregate human preferences but often fails to address the diverse and confli…
cs.RO2025
LLaMAR: Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments
Siddharth Nayak, Adelmo Morrison Orozco, Marina Ten Have +10
The ability of Language Models (LMs) to understand natural language makes them a powerful tool for parsing human instructions into task plans for autonomous robots. Unlike traditio…