most citedThe Bitter Lesson Learned from 2,000+ Multilingual Benchmarks

1 citations · 1 across the 3 of their papers we have counts for

collaborators

8 papers

cs.CL2025

Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translation

Sherrie Shen, Weixuan Wang, Alexandra Birch

The faithful transfer of contextually-embedded meaning continues to challenge contemporary machine translation (MT), particularly in the rendering of culture-bound terms--expressio…

cs.CL2025

Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization

Weixuan Wang, Minghao Wu, Barry Haddow +1

Long document summarization remains a significant challenge for current large language models (LLMs), as existing approaches commonly struggle with information loss, factual incons…

cs.CL2025

ExpertSteer: Intervening in LLMs through Expert Knowledge

Weixuan Wang, Minghao Wu, Barry Haddow +1

Large Language Models (LLMs) exhibit remarkable capabilities across various tasks, yet guiding them to follow desired behaviours during inference remains a significant challenge. A…

cs.CL2025

HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language Models

Weixuan Wang, Minghao Wu, Barry Haddow +1

Fine-tuning large language models (LLMs) on a mixture of diverse datasets poses challenges due to data imbalance and heterogeneity. Existing methods often address these issues acro…

cs.CL20251 cited

The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks

Minghao Wu, Weixuan Wang, Sinuo Liu +7

As large language models (LLMs) continue to advance in linguistic capabilities, robust multilingual evaluation has become essential for promoting equitable technological progress.…

cs.SE2025

PredicateFix: Repairing Static Analysis Alerts with Bridging Predicates

Yuan-An Xiao, Weixuan Wang, Dong Liu +3

Fixing static analysis alerts in source code with Large Language Models (LLMs) is becoming increasingly popular. However, LLMs often hallucinate and perform poorly for complex and…