2 papers
cs.CL2026
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham +3
Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduc…
cs.CL2026
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham +2
Recently, Large Language Models (LLMs) have gained significant traction in medical domain, especially in developing a QA systems to Medical QA systems for enhancing access to healt…