works on

From the 1 of 11 linked papers with an AI index.

collaborators

11 papers

cs.CL2026

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

Sushant Gautam, Vajira Thambawita, Michael A. Riegler +2

The paper studies how design choices affect the reliability and interpretability of multimodal visual question answering systems for gastrointestinal endoscopy, finding that struct…

cs.CL2026

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

Oriana Presacan, Andreea Grama, Larisa Irimină +6

Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insu…

cs.LG2026

Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis

Michael A. Riegler, Birk Sebastian Frostelid Torpmann-Hagen

The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single featur…

cs.CV2026

VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations

Sushant Gautam, Cise Midoglu, Vajira Thambawita +2

Hallucinations in video-capable vision-language models (Video-VLMs) remain frequent and high-confidence, while existing uncertainty metrics often fail to align with correctness. We…

cs.CV2025

HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models

Sushant Gautam, Michael A. Riegler, PÃ¥l Halvorsen

Vision-language models (VLMs) enable open-ended visual question answering but remain prone to hallucinations. We present HEDGE, a unified framework for hallucination detection that…

stat.ME2025

Using Large Language Models to Suggest Informative Prior Distributions in Bayesian Statistics

Michael A. Riegler, Kristoffer Herland Hellton, Vajira Thambawita +1

Selecting prior distributions in Bayesian statistics is challenging, resource-intensive, and subjective. We analyze using large-language models (LLMs) to suggest suitable, knowledg…