Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
Mason Nakamura, Saaduddin Mahmud, Kyle H. Wray +2
Aligning LLMs with user preferences is crucial for real-world use but often requires costly fine-tuning or expensive inference, forcing trade-offs between alignment quality and com…
cs.LG2024
MAPLE: A Framework for Active Preference Learning Guided by Large Language Models
Saaduddin Mahmud, Mason Nakamura, Shlomo Zilberstein
The advent of large language models (LLMs) has sparked significant interest in using natural language for preference learning. However, existing methods often suffer from high comp…