7 papers
SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization
Usman Naseem, Robert Geislinger, Juan Ren +31
We present SemEval-2026 Task 9, a shared task on online polarization detection, covering 22 languages and comprising over 110K annotated instances. Each data instance is multi-labe…
Scaling Truth: The Confidence Paradox in AI Fact-Checking
Ihsan A. Qazi, Zohaib Khan, Abdullah Ghani +7
The rise of misinformation underscores the need for scalable and reliable fact-checking solutions. Large language models (LLMs) hold promise in automating fact verification, yet th…
TweakLLM: A Routing Architecture for Dynamic Tailoring of Cached Responses
Muhammad Taha Cheema, Abeer Aamir, Khawaja Gul Muhammad +3
Large Language Models (LLMs) process millions of queries daily, making efficient response caching a compelling optimization for reducing cost and latency. However, preserving relev…
Semantic Caching for Improving Web Affordability
Hafsa Akbar, Danish Athar, Muhammad Ayain Fida Rana +4
The rapid growth of web content has led to increasingly large webpages, posing significant challenges for Internet affordability, especially in developing countries where data cost…
To Label or Not to Label: Hybrid Active Learning for Neural Machine Translation
Abdul Hameed Azeemi, Ihsan Ayyub Qazi, Agha Ali Raza
Active learning (AL) techniques reduce labeling costs for training neural machine translation (NMT) models by selecting smaller representative subsets from unlabeled data for annot…
Language Model-Driven Data Pruning Enables Efficient Active Learning
Abdul Hameed Azeemi, Ihsan Ayyub Qazi, Agha Ali Raza
Active learning (AL) optimizes data labeling efficiency by selecting the most informative instances for annotation. A key component in this procedure is an acquisition function tha…