7 papers
Quantifying and Mitigating Self-Preference Bias of LLM Judges
Jinming Yang, Zheng Hu, Chuxian Qiu +3
LLM-as-a-Judge has become a dominant approach in automated evaluation systems, playing critical roles in model alignment, leaderboard construction, quality control, and so on. Howe…
The Automatic Verification of Image-Text Claims (AVerImaTeC) Shared Task
Rui Cao, Zhenyun Deng, Yulong Chen +2
The Automatic Verification of Image-Text Claims (AVerImaTeC) shared task aims to advance system development for retrieving evidence and verifying real-world image-text claims. Part…
Multimodal Claim Extraction for Fact-Checking
Joycelyn Teo, Rui Cao, Zhenyun Deng +3
Automated Fact-Checking (AFC) relies on claim extraction as a first step, yet existing methods largely overlook the multimodal nature of today's misinformation. Social media posts…
Improving Zero-shot Sentence Decontextualisation with Content Selection and Planning
Zhenyun Deng, Yulong Chen, Andreas Vlachos
Extracting individual sentences from a document as evidence or reasoning steps is commonly done in many NLP tasks. However, extracted sentences often lack context necessary to make…
PledgeTracker: A System for Monitoring the Fulfilment of Pledges
Yulong Chen, Michael Sejr Schlichtkrull, Zhenyun Deng +5
Political pledges reflect candidates' policy commitments, but tracking their fulfilment requires reasoning over incremental evidence distributed across multiple, dynamically update…
Hard negative sampling in hyperedge prediction
Zhenyu Deng, Tao Zhou, Yilin Bi
Hypergraph, which allows each hyperedge to encompass an arbitrary number of nodes, is a powerful tool for modeling multi-entity interactions. Hyperedge prediction is a fundamental…