5 papers
Scaling Truth: The Confidence Paradox in AI Fact-Checking
Ihsan A. Qazi, Zohaib Khan, Abdullah Ghani +7
The rise of misinformation underscores the need for scalable and reliable fact-checking solutions. Large language models (LLMs) hold promise in automating fact verification, yet th…
The Fellowship of the LLMs: Multi-Model Workflows for Synthetic Preference Optimization Dataset Generation
Samee Arif, Sualeha Farid, Abdul Hameed Azeemi +2
This paper presents a novel methodology for generating synthetic Preference Optimization (PO) datasets using multi-model workflows. We evaluate the effectiveness and potential of t…
To Label or Not to Label: Hybrid Active Learning for Neural Machine Translation
Abdul Hameed Azeemi, Ihsan Ayyub Qazi, Agha Ali Raza
Active learning (AL) techniques reduce labeling costs for training neural machine translation (NMT) models by selecting smaller representative subsets from unlabeled data for annot…
Language Model-Driven Data Pruning Enables Efficient Active Learning
Abdul Hameed Azeemi, Ihsan Ayyub Qazi, Agha Ali Raza
Active learning (AL) optimizes data labeling efficiency by selecting the most informative instances for annotation. A key component in this procedure is an acquisition function tha…
Generalists vs. Specialists: Evaluating Large Language Models for Urdu
Samee Arif, Abdul Hameed Azeemi, Agha Ali Raza +1
In this paper, we compare general-purpose models, GPT-4-Turbo and Llama-3-8b, with special-purpose models--XLM-Roberta-large, mT5-large, and Llama-3-8b--that have been fine-tuned o…