activity
20202026
most citedThe SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods

29 citations · 30 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

Bowen Wang, Youwen Zhang, Ritesh Mehta

We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection…

cs.CV2025

PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning

Jiahao Zhang, Bowen Wang, Hong Liu +2

Visual In-Context Learning (VICL) uses input-output image pairs, referred to as in-context pairs (or examples), as prompts alongside query images to guide models in performing dive…

cs.CV2025

E-InMeMo: Enhanced Prompting for Visual In-Context Learning

Jiahao Zhang, Bowen Wang, Hong Liu +3

Large-scale models trained on extensive datasets have become the standard due to their strong generalizability across diverse tasks. In-context learning (ICL), widely used in natur…

cs.CV2024

ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training

Zhouqiang Jiang, Bowen Wang, Junhao Chen +1

Recent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obvi…

cs.CV2024

Explainable Image Recognition via Enhanced Slot-attention Based Classifier

Bowen Wang, Liangzhi Li, Jiahao Zhang +2

The imperative to comprehend the behaviors of deep learning models is of utmost importance. In this realm, Explainable Artificial Intelligence (XAI) has emerged as a promising aven…

cs.CV20231 cited

Instruct Me More! Random Prompting for Visual In-Context Learning

Jiahao Zhang, Bowen Wang, Liangzhi Li +2

Large-scale models trained on extensive datasets, have emerged as the preferred approach due to their high generalizability across various tasks. In-context learning (ICL), a popul…