5 papers
Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs
Guiqiu Liao, Matjaz Jogan, Daniel A. Hashimoto
Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leading to costly inference and lim…
Active Learning for Efficient Annotation of Surgical Videos with Weak Supervision
Manasa Dendukuri, Matjaz Jogan, Daniel A. Hashimoto +1
Precise spatial-temporal annotation of laparoscopic videos is time-consuming and requires expert knowledge. We propose a human-in-the-loop knowledge acquisition framework that comb…
Future Slot Prediction for Unsupervised Object Discovery in Surgical Video
Guiqiu Liao, Matjaz Jogan, Marcel Hussing +3
Object-centric slot attention is an emerging paradigm for unsupervised learning of structured, interpretable object-centric representations (slots). This enables effective reasonin…
FORLA: Federated Object-centric Representation Learning with Slot Attention
Guiqiu Liao, Matjaz Jogan, Eric Eaton +1
Learning efficient visual representations across heterogeneous unlabeled datasets remains a central challenge in federated learning. Effective federated representations require fea…
Slot-BERT: Self-supervised Object Discovery in Surgical Video
Guiqiu Liao, Matjaz Jogan, Marcel Hussing +5
Object-centric slot attention is a powerful framework for unsupervised learning of structured and explainable representations that can support reasoning about objects and actions,…