most citedWhat a Whole Slide Image Can Tell? Subtype-guided Masked Transformer for Pathological Image Captioning

3 citations · 7 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20243 cited

PathInsight: Instruction Tuning of Multimodal Datasets and Models for Intelligence Assisted Diagnosis in Histopathology

Xiaomin Wu, Rui Xu, Pengchen Wei +4

Pathological diagnosis remains the definitive standard for identifying tumors. The rise of multimodal large models has simplified the process of integrating image analysis with tex…

cs.CV20231 cited

SCAAT: Improving Neural Network Interpretability via Saliency Constrained Adaptive Adversarial Training

Rui Xu, Wenkang Qin, Peixiang Huang +2

Deep Neural Networks (DNNs) are expected to provide explanation for users to understand their black-box predictions. Saliency map is a common form of explanation illustrating the h…

cs.CV2023

Improving Vision-and-Language Reasoning via Spatial Relations Modeling

Cheng Yang, Rui Xu, Ye Guo +5

Visual commonsense reasoning (VCR) is a challenging multi-modal task, which requires high-level cognition and commonsense reasoning ability about the real world. In recent years, l…

cs.CV20233 cited

What a Whole Slide Image Can Tell? Subtype-guided Masked Transformer for Pathological Image Captioning

Wenkang Qin, Rui Xu, Peixiang Huang +3

Pathological captioning of Whole Slide Images (WSIs), though is essential in computer-aided pathological diagnosis, has rarely been studied due to the limitations in datasets and m…

eess.IV2023

Assessing and Enhancing Robustness of Deep Learning Models with Corruption Emulation in Digital Pathology

Peixiang Huang, Songtao Zhang, Yulu Gan +6

Deep learning in digital pathology brings intelligence and automation as substantial enhancements to pathological analysis, the gold standard of clinical diagnosis. However, multip…