5 papers
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output
Guozheng Li, Xiyan Fu, Yiwen Guo
Current reinforcement learning from human feedback (RLHF) methods primarily rely on scalar rewards from a trained reward model (RM). While effective, scalar rewards are often noisy…
When Languages Disagree: Self-Evolving Multilingual LLM Judges
Xiyan Fu, Wei Lu
Multilingual LLM-as-a-judge is widely used to evaluate model outputs across languages, but suffers from cross-lingual inconsistency (Fu and Liu, 2025). Existing methods typically t…
Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization
Xiyan Fu, Wei Liu
Compositional generalization refers to correctly interpret novel combinations of known primitives, which remains a major challenge. Existing approaches often rely on supervised fin…
How Reliable is Multilingual LLM-as-a-Judge?
Xiyan Fu, Wei Liu
LLM-as-a-Judge has emerged as a popular evaluation strategy, where advanced large language models assess generation results in alignment with human instructions. While these models…
The Mystery of Compositional Generalization in Graph-based Generative Commonsense Reasoning
Xiyan Fu, Anette Frank
While LLMs have emerged as performant architectures for reasoning tasks, their compositional generalization capabilities have been questioned. In this work, we introduce a Composit…