activity
20242026
collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

MedSNIP: Building and Benchmarking Snippet-Level Granularity for Medical Fact Verification

Hasan Iqbal, Sarfraz Ahmad, Hyunjae Kim +5

A medical claim's correctness often depends not on the claim alone, but on the clinical structure around it. A claim may require a lab reference range, a causal or conditional link…

cs.CL2026

Jais 2: A Family of Arabic-Centric Open Large Language Models

Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim +57

Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong p…

cs.CL2026

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

Ahmer Tabassum, Sarfraz Ahmad, Hasan Iqbal +3

Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark…

cs.CL2025

FRaN-X: FRaming and Narratives-eXplorer

Artur Muratov, Hana Fatima Shaikh, Vanshikaa Jani +21

We present FRaN-X, a Framing and Narratives Explorer that automatically detects entity mentions and classifies their narrative roles directly from raw text. FRaN-X comprises a two-…

cs.CL2025

NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors

Numaan Naeem, Sarfraz Ahmad, Momina Ahsan +1

This paper presents our system for Track 1: Mistake Identification in the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors. The task involves evaluating…

cs.CL2025

UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking

Sarfraz Ahmad, Hasan Iqbal, Momina Ahsan +6

The rapid adoption of Large Language Models (LLMs) has raised important concerns about the factual reliability of their outputs, particularly in low-resource languages such as Urdu…