From the 1 of 6 linked papers with an AI index.
6 papers
Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs
Guiqiu Liao, Matjaz Jogan, Daniel A. Hashimoto
Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leading to costly inference and lim…
Active Learning for Efficient Annotation of Surgical Videos with Weak Supervision
Manasa Dendukuri, Matjaz Jogan, Daniel A. Hashimoto +1
The paper presents a human‑in‑the‑loop framework that combines active learning with dual‑loss weak supervision to cut the effort needed for annotating laparoscopic video frames, en…
Slot-BERT: Self-supervised Object Discovery in Surgical Video
Guiqiu Liao, Matjaz Jogan, Marcel Hussing +5
Object-centric slot attention is a powerful framework for unsupervised learning of structured and explainable representations that can support reasoning about objects and actions,…
FORLA: Federated Object-centric Representation Learning with Slot Attention
Guiqiu Liao, Matjaz Jogan, Eric Eaton +1
Learning efficient visual representations across heterogeneous unlabeled datasets remains a central challenge in federated learning. Effective federated representations require fea…
Future Slot Prediction for Unsupervised Object Discovery in Surgical Video
Guiqiu Liao, Matjaz Jogan, Marcel Hussing +3
Object-centric slot attention is an emerging paradigm for unsupervised learning of structured, interpretable object-centric representations (slots). This enables effective reasonin…
Disentangling spatio-temporal knowledge for weakly supervised object detection and segmentation in surgical video
Guiqiu Liao, Matjaz Jogan, Sai Koushik +2
Weakly supervised video object segmentation (WSVOS) enables the identification of segmentation maps without requiring an extensive training dataset of object masks, relying instead…