collaborators

8 papers

cs.CL2026

Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad +1

Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems. As Large…

cs.CL2026

Riazi-8B: An Urdu Large Language Model for Mathematical Reasoning

Azher Ali, Ibtsam Haider, Raja Khurram Shahzad +2

Recent LLMs demonstrate strong mathematical reasoning capabilities, but existing gains rely heavily on English-centric training resources and benchmarks. As a result, reasoning per…

cs.CL2026

ROMEVA: Geometry-Preserving Vocabulary Expansion for Roman Urdu Language Models

Mahnoor Khan, Afsheen Asif, Milhan Afzal Khan +2

Multilingual Language Models like mBERT are widely used for low-resource NLP, yet their adaptation to morphologically inconsistent languages such as Roman Urdu remains underexplore…

cs.CL2026

Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form Generation

Areeba Hassan, Arooj Kausar, Syeda Kisaa Fatima +2

Large Language Models (LLMs) generate fluent long-form text, however, often add unsupported factual claims. Existing verification techniques improve factuality by grounding generat…

cs.AI2026

PRIME: Evaluating Prompt Resolution Under Incompatible Instructions in LLMs

Tehreem Javed, Shumaim Fatimah, Masooma Bakhtiari +2

Large language models (LLMs) often encounter conflicting prompts, although current instruction following benchmarks assess those meta-instructions in isolation, limiting the insigh…

cs.AI2026

From Affect Prediction to Affect Forecasting: Evidence for Distinct Information Sources in Longitudinal Text

Sadia Noor, Seemab Latif, Raja Khurram Shahzad +1

Modeling dimensional affect in longitudinal text requires distinguishing current affect estimation from future affective change forecasting. Existing approaches often treat each te…