works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CL2026

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

Sushant Gautam, Vajira Thambawita, Michael A. Riegler +2

The paper studies how design choices affect the reliability and interpretability of multimodal visual question answering systems for gastrointestinal endoscopy, finding that struct…

cs.LG2026

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels

Sushant Gautam, Finn Schwall, Annika Willoch Olstad +6

Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. We formalize this sett…

cs.SI2026

The Moltbook Observatory Archive: an incremental dataset of agent-only social network activity

Sushant Gautam, Annika W. Olstad, Klas H. Pettersen +1

Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory Archive, an incremental datase…

cs.CV2026

VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations

Sushant Gautam, Cise Midoglu, Vajira Thambawita +2

Hallucinations in video-capable vision-language models (Video-VLMs) remain frequent and high-confidence, while existing uncertainty metrics often fail to align with correctness. We…

cs.CV2025

HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models

Sushant Gautam, Michael A. Riegler, PÃ¥l Halvorsen

Vision-language models (VLMs) enable open-ended visual question answering but remain prone to hallucinations. We present HEDGE, a unified framework for hallucination detection that…

cs.CV2025

Medico 2025: Visual Question Answering for Gastrointestinal Imaging

Sushant Gautam, Vajira Thambawita, Michael Riegler +2

The Medico 2025 challenge addresses Visual Question Answering (VQA) for Gastrointestinal (GI) imaging, organized as part of the MediaEval task series. The challenge focuses on deve…