4 papers
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham +3
Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduc…
Progressive Code Integration for Abstractive Bug Report Summarization
Shaira Sadia Karim, Abrar Mahmud Rahim, Lamia Alam +4
Bug reports are often unstructured and verbose, making it challenging for developers to efficiently comprehend software issues. Existing summarization approaches typically rely on…
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
Ishmam Tashdeed, Md. Atiqur Rahman, Sabrina Islam +1
Personalized federated learning (PFL) possesses the unique capability of preserving data confidentiality among clients while tackling the data heterogeneity problem of non-independ…
Visual Robustness Benchmark for Visual Question Answering (VQA)
Md Farhan Ishmam, Ishmam Tashdeed, Talukder Asir Saadat +3
Can Visual Question Answering (VQA) systems perform just as well when deployed in the real world? Or are they susceptible to realistic corruption effects e.g. image blur, which can…