activity
20212026
most citedCD-Net: Histopathology Representation Learning using Pyramidal Context-Detail Network

3 citations · 7 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

Yumi Lee, Harim Oh, Hyoryung Kim +52

The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remai…

cs.CV2026

Semantic Context-aware mOdality fUsion Transformer (SCOUT): A Context-Aware Multimodal Transformer for Concept-Grounded Pathology Report Generation

Suryakant Singh, Saarthak Kapse, Joel Saltz +1

Whole-slide pathology report generation requires models to integrate localized histomorphology, global tissue context, and diagnostically relevant semantic information, yet existin…

cs.CV2025

TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning

Varun Belagali, Saarthak Kapse, Pierre Marza +12

The interpretation of small tiles in large whole slide images (WSI) often needs a larger image context. We introduce TICON, a transformer-based tile representation contextualizer t…

cs.CV2025

PEaRL: Pathway-Enhanced Representation Learning for Gene and Pathway Expression Prediction from Histology

Sejuti Majumder, Saarthak Kapse, Moinak Bhattacharya +3

Integrating histopathology with spatial transcriptomics (ST) provides a powerful opportunity to link tissue morphology with molecular function. Yet most existing multimodal approac…

cs.CV2025

GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology

Saarthak Kapse, Pushpak Pati, Srikar Yellapragada +5

Pretraining a Multiple Instance Learning (MIL) aggregator enables the derivation of Whole Slide Image (WSI)-level embeddings from patch-level representations without supervision. W…

cs.CV2025

Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing

Saarthak Kapse, Robin Betz, Srinivasan Sivanandan

State Space Models (SSMs) with selective scan (Mamba) have been adapted into efficient vision models. Mamba, unlike Vision Transformers, achieves linear complexity for token intera…