1 citations · 1 across the 5 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
Yushi Hu, Reyhane Askari-Hemmat, Melissa Hall +3
Reward models (RMs) are essential for training large language models (LLMs), but remain underexplored for omni models that handle interleaved image and text sequences. We introduce…
cs.CL2024★ 1 cited
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
Anaelia Ovalle, Krunoslav Lehman Pavasovic, Louis Martin +5
Natural-language assistants are designed to provide users with helpful responses while avoiding harmful outputs, largely achieved through alignment to human preferences. Yet there…