1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Tao Zhang, Xiangtai Li, Zilong Huang +6
Multimodal Large Language Models (MLLMs) achieve remarkable performance for fine-grained pixel-level understanding tasks. However, all the works rely heavily on extra components, s…
cs.CV2024★ 1 cited
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Kevin Qinghong Lin, Linjie Li, Difei Gao +6
Building Graphical User Interface (GUI) assistants holds significant promise for enhancing human workflow productivity. While most agents are language-based, relying on closed-sour…
cs.CV2022★ 1 cited
PCCT: Progressive Class-Center Triplet Loss for Imbalanced Medical Image Classification
Kanghao Chen, Weixian Lei, Rong Zhang +3
Imbalanced training data is a significant challenge for medical image classification. In this study, we propose a novel Progressive Class-Center Triplet (PCCT) framework to allevia…