8 papers
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
Firoj Alam, Gagan Bhatia, Sahinur Rahman Laskar +1
While Large Language Models (LLMs) are increasingly adopted as automated judges for evaluating generated text, their outputs are often costly, and highly sensitive to prompt design…
CoCR-RAG: Enhancing Retrieval-Augmented Generation in Web Q&A via Concept-oriented Context Reconstruction
Kaize Shi, Xueyao Sun, Qika Lin +4
Retrieval-augmented generation (RAG) has shown promising results in enhancing Q&A by incorporating information from the web and other external sources. However, the supporting docu…
MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi +2
Audio large language models (AudioLLMs) enable instruction-following over speech and general audio, but progress is increasingly limited by the lack of diverse, conversational, ins…
Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models
Basel Mousi, Fahim Dalvi, Shammur Chowdhury +2
Vision-language models (VLMs) can achieve high accuracy while still accepting culturally plausible but visually incorrect interpretations. Existing hallucination benchmarks rarely…
Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
Hunzalah Hassan Bhatti, Firoj Alam
Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal content remains uneven across languages.…
The Landscape of Arabic Large Language Models (ALLMs): A New Era for Arabic Language Technology
Shahad Al-Khalifa, Nadir Durrani, Hend Al-Khalifa +1
The emergence of ChatGPT marked a transformative milestone for Artificial Intelligence (AI), showcasing the remarkable potential of Large Language Models (LLMs) to generate human-l…