From the 1 of 11 linked papers with an AI index.
11 papers
Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA
Sushant Gautam, Vajira Thambawita, Michael A. Riegler +2
The paper studies how design choices affect the reliability and interpretability of multimodal visual question answering systems for gastrointestinal endoscopy, finding that struct…
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
Oriana Presacan, Andreea Grama, Larisa IriminÄ +6
Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insu…
Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis
Michael A. Riegler, Birk Sebastian Frostelid Torpmann-Hagen
The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single featur…
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
Sushant Gautam, Cise Midoglu, Vajira Thambawita +2
Hallucinations in video-capable vision-language models (Video-VLMs) remain frequent and high-confidence, while existing uncertainty metrics often fail to align with correctness. We…
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
Sushant Gautam, Michael A. Riegler, PÃ¥l Halvorsen
Vision-language models (VLMs) enable open-ended visual question answering but remain prone to hallucinations. We present HEDGE, a unified framework for hallucination detection that…
Using Large Language Models to Suggest Informative Prior Distributions in Bayesian Statistics
Michael A. Riegler, Kristoffer Herland Hellton, Vajira Thambawita +1
Selecting prior distributions in Bayesian statistics is challenging, resource-intensive, and subjective. We analyze using large-language models (LLMs) to suggest suitable, knowledg…