Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
Jingyan Shen, Jiarui Yao, Rui Yang +5
Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, rew…
cs.AI2024
On the Expressive Power of Tree-Structured Probabilistic Circuits
Lang Yin, Han Zhao
Probabilistic circuits (PCs) have emerged as a powerful framework to compactly represent probability distributions for efficient and exact probabilistic inference. It has been show…