12 papers
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
Gagan Bhatia, Julian Schlenker, Simone Paolo Ponzetto +1
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and the…
From RAG to Agentic RAG for Faithful Islamic Question Answering
Gagan Bhatia, Hamdy Mubarak, Mustafa Jarrar +8
Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. Yet standard MCQ/MRC-style e…
Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA
Ummar Abbas, Mourad Ouzzani, Mohamed Y. Eltabakh +7
Large language models (LLMs) can answer religious knowledge queries fluently, yet they often hallucinate and misattribute sources, which is especially consequential in Islamic sett…
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli +10
Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and ho…
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
Subhadeep Roy, Gagan Bhatia, Steffen Eger
Automatic metrics are widely used to evaluate text-to-image models, often replacing human judgment in benchmarking, model selection, and large-scale data filtering. Yet they may re…
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
Firoj Alam, Gagan Bhatia, Sahinur Rahman Laskar +1
While Large Language Models (LLMs) are increasingly adopted as automated judges for evaluating generated text, their outputs are often costly, and highly sensitive to prompt design…