most citedLocalizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs

3 citations · 3 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2026

Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

Ta Duc Huy, Trang Nguyen, Townim Chowdhury +5

Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rel…

cs.CV20263 cited

Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs

Dung Nguyen, Minh Khoi Ho, Huy Ta +11

Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradic…

cs.SD2026

AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs

Townim Faisal Chowdhury, Ta Duc Huy, Siqi Pan +2

Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability i…

cs.CV2025

From Healthy Scans to Annotated Tumors: A Tumor Fabrication Framework for 3D Brain MRI Synthesis

Nayu Dong, Townim Chowdhury, Hieu Phan +3

The scarcity of annotated Magnetic Resonance Imaging (MRI) tumor data presents a major obstacle to accurate and automated tumor segmentation. While existing data synthesis methods…

cs.CV2025

Looking in the mirror: A faithful counterfactual explanation method for interpreting deep image classification models

Townim Faisal Chowdhury, Vu Minh Hieu Phan, Kewen Liao +5

Counterfactual explanations (CFE) for deep image classifiers aim to reveal how minimal input changes lead to different model decisions, providing critical insights for model interp…