10 papers
Zero-Shot Captioning for Cultural Heritage: Automated Image Analysis of Traditional Indonesian Clothing
Anugrah Aidin Yotolembah, Novanto Yudistira, Gembong Edhi Setyawan
This paper presents Custom ZeroCLIP, a retrieval-augmented vision-language framework for zero-shot captioning of Indonesian traditional garments. The dataset contains 3,800 expert-…
Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study
Pieter Christy Yan Yudhistira, Dzaki Rafif Malik, Novanto Yudistira
Medical Vision-Language Models (VLMs) are typically evaluated on English radiology visual question answering benchmarks, leaving their robustness under non-English clinical languag…
Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification
One Octadion, Novanto Yudistira, Lailil Muflikhah
Deep learning models for medical imaging often exhibit overconfidence, creating safety risks in ambiguous diagnostic scenarios. While Conformal Prediction (CP) provides distributio…
Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints
Mutiara Shabrina, Nova Kurnia Putri, Jefri Satria Ferdiansyah +2
Text-driven image manipulation often suffers from attribute entanglement, where modifying a target attribute (e.g., adding bangs) unintentionally alters other semantic properties s…
Input-Adaptive Visual Preprocessing for Efficient Fast Vision-Language Model Inference
Putu Indah Githa Cahyani, Komang David Dananjaya Suartana, Novanto Yudistira
Vision-Language Models (VLMs) have demonstrated strong performance on multimodal reasoning tasks, but their deployment remains challenging due to high inference latency and computa…
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
Novanto Yudistira
This study introduces a pioneering methodology for human action recognition by harnessing deep neural network techniques and adaptive fusion strategies across multiple modalities,…