4 papers
SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics
Silvia Casola, Ryan Soh-Eun Shim, Felicia Körner +2
An increasing body of work has leveraged multilingual language models for Natural Language Generation tasks such as summarization. A major empirical bottleneck in this area is the…
LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task
Elisa Leonardelli, Silvia Casola, Siyao Peng +8
Many researchers have reached the conclusion that AI models should be trained to be aware of the possibility of variation and disagreement in human judgments, and evaluated as per…
Reason to Rote: Rethinking Memorization in Reasoning
Yupei Du, Philipp Mondorf, Silvia Casola +3
Large language models readily memorize arbitrary training instances, such as label noise, yet they perform strikingly well on reasoning tasks. In this work, we investigate how lang…
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
Silvia Casola, Yang Janet Liu, Siyao Peng +3
Human language production exhibits remarkable richness and variation, reflecting diverse communication styles and intents. However, this variation is often overlooked in summarizat…