From the 1 of 9 linked papers with an AI index.
9 papers
Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA
Sushant Gautam, Vajira Thambawita, Michael A. Riegler +2
The paper studies how design choices affect the reliability and interpretability of multimodal visual question answering systems for gastrointestinal endoscopy, finding that struct…
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
Sushant Gautam, Finn Schwall, Annika Willoch Olstad +6
Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. We formalize this sett…
The Moltbook Observatory Archive: an incremental dataset of agent-only social network activity
Sushant Gautam, Annika W. Olstad, Klas H. Pettersen +1
Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory Archive, an incremental datase…
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
Sushant Gautam, Cise Midoglu, Vajira Thambawita +2
Hallucinations in video-capable vision-language models (Video-VLMs) remain frequent and high-confidence, while existing uncertainty metrics often fail to align with correctness. We…
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
Sushant Gautam, Michael A. Riegler, PÃ¥l Halvorsen
Vision-language models (VLMs) enable open-ended visual question answering but remain prone to hallucinations. We present HEDGE, a unified framework for hallucination detection that…
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
Sushant Gautam, Vajira Thambawita, Michael Riegler +2
The Medico 2025 challenge addresses Visual Question Answering (VQA) for Gastrointestinal (GI) imaging, organized as part of the MediaEval task series. The challenge focuses on deve…