4 papers · 1 filter
When Languages Disagree: Self-Evolving Multilingual LLM Judges
Xiyan Fu, Wei Lu
Multilingual LLM-as-a-judge is widely used to evaluate model outputs across languages, but suffers from cross-lingual inconsistency (Fu and Liu, 2025). Existing methods typically t…
How Reliable is Multilingual LLM-as-a-Judge?
Xiyan Fu, Wei Liu
LLM-as-a-Judge has emerged as a popular evaluation strategy, where advanced large language models assess generation results in alignment with human instructions. While these models…
The Mystery of Compositional Generalization in Graph-based Generative Commonsense Reasoning
Xiyan Fu, Anette Frank
While LLMs have emerged as performant architectures for reasoning tasks, their compositional generalization capabilities have been questioned. In this work, we introduce a Composit…
Exploring Continual Learning of Compositional Generalization in NLI
Xiyan Fu, Anette Frank
Compositional Natural Language Inference has been explored to assess the true abilities of neural models to perform NLI. Yet, current evaluations assume models to have full access…