8 papers
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad +1
Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems. As Large…
Riazi-8B: An Urdu Large Language Model for Mathematical Reasoning
Azher Ali, Ibtsam Haider, Raja Khurram Shahzad +2
Recent LLMs demonstrate strong mathematical reasoning capabilities, but existing gains rely heavily on English-centric training resources and benchmarks. As a result, reasoning per…
ROMEVA: Geometry-Preserving Vocabulary Expansion for Roman Urdu Language Models
Mahnoor Khan, Afsheen Asif, Milhan Afzal Khan +2
Multilingual Language Models like mBERT are widely used for low-resource NLP, yet their adaptation to morphologically inconsistent languages such as Roman Urdu remains underexplore…
Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form Generation
Areeba Hassan, Arooj Kausar, Syeda Kisaa Fatima +2
Large Language Models (LLMs) generate fluent long-form text, however, often add unsupported factual claims. Existing verification techniques improve factuality by grounding generat…
PRIME: Evaluating Prompt Resolution Under Incompatible Instructions in LLMs
Tehreem Javed, Shumaim Fatimah, Masooma Bakhtiari +2
Large language models (LLMs) often encounter conflicting prompts, although current instruction following benchmarks assess those meta-instructions in isolation, limiting the insigh…
From Affect Prediction to Affect Forecasting: Evidence for Distinct Information Sources in Longitudinal Text
Sadia Noor, Seemab Latif, Raja Khurram Shahzad +1
Modeling dimensional affect in longitudinal text requires distinguishing current affect estimation from future affective change forecasting. Existing approaches often treat each te…