5 papers
Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Quality Evaluation
Weilu Xu, Yunzhi Shen, Xinye Wang +2
Large language models (LLMs) have shown strong potential in fine-grained translation quality evaluation (QE), yet existing MQM-based approaches typically rely on fixed rubric confi…
Unlocking Fine-Grained Translation Quality Estimation in LRMs through Synergistically Evolving Implicit and Explicit Reasoning
Renfei Dang, Xinye Wang, Zhejian Lai +5
Large Reasoning Models (LRMs) still struggle with fine-grained translation quality estimation (QE), even with long reasoning chains. We argue that LRMs already possess strong multi…
Scale Determines Whether Language Models Organize Representation Geometry for Prediction
Weilun Xu
In language models, what a representation encodes is determined by the geometry of its representation space: distances, not activations, carry meaning. Existing tools characterize…
Diachronic Modeling of Tonal Coherence on the Tonnetz Across Classical and Popular Repertoires
Weilun Xu, Edward Hall, Martin Rohrmeier
How do different musical traditions achieve tonal coherence? Most computational measures to date have analysed tonal coherence in terms of a single dimension, whereas a multi-dimen…
Probing Ethical Framework Representations in Large Language Models: Structure, Entanglement, and Methodological Challenges
Weilun Xu, Alexander Rusnak, Frederic Kaplan
When large language models make ethical judgments, do their internal representations distinguish between normative frameworks, or collapse ethics into a single acceptability dimens…