works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs

Guiqiu Liao, Matjaz Jogan, Daniel A. Hashimoto

Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leading to costly inference and lim…

cs.CV2026

Active Learning for Efficient Annotation of Surgical Videos with Weak Supervision

Manasa Dendukuri, Matjaz Jogan, Daniel A. Hashimoto +1

The paper presents a human‑in‑the‑loop framework that combines active learning with dual‑loss weak supervision to cut the effort needed for annotating laparoscopic video frames, en…

cs.CV2026

DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction

Guiqiu Liao, Matjaž Jogan, Daniel A. Hashimoto

Dense prediction tasks in surgical computer vision, such as segmentation and surgical zone prediction, can provide valuable guidance for laparoscopic and robotic surgery. However,…

cs.CV2025

FORLA: Federated Object-centric Representation Learning with Slot Attention

Guiqiu Liao, Matjaz Jogan, Eric Eaton +1

Learning efficient visual representations across heterogeneous unlabeled datasets remains a central challenge in federated learning. Effective federated representations require fea…

cs.CV2025

Future Slot Prediction for Unsupervised Object Discovery in Surgical Video

Guiqiu Liao, Matjaz Jogan, Marcel Hussing +3

Object-centric slot attention is an emerging paradigm for unsupervised learning of structured, interpretable object-centric representations (slots). This enables effective reasonin…

cs.CV2024

Disentangling spatio-temporal knowledge for weakly supervised object detection and segmentation in surgical video

Guiqiu Liao, Matjaz Jogan, Sai Koushik +2

Weakly supervised video object segmentation (WSVOS) enables the identification of segmentation maps without requiring an extensive training dataset of object masks, relying instead…