1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.LG2026★ 1 cited
One Token to Fool LLM-as-a-Judge
Yulai Zhao, Haolin Liu, Dian Yu +4
Large language models (LLMs) are increasingly trusted as automated judges, assisting evaluation and providing reward signals for training other models, particularly in reference-ba…
cs.LG2025
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
Dian Yu, Yulai Zhao, Kishan Panaganti +3
We propose Reinforcement Learning with Explicit Human Values (RLEV), a method that aligns Large Language Model (LLM) optimization directly with quantifiable human value signals. Wh…