activity
20212026
most citedLLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

231 citations · 504 across the 19 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao +27

Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and treatment selection by learning trans…

cs.CV2026

Learning Sparse Visual Representations via Spatial-Semantic Factorization

Theodore Zhengde Zhao, Sid Kiblawi, Jianwei Yang +6

Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens th…

cs.CV2025

Boltzmann Attention Sampling for Image Analysis with Small Objects

Theodore Zhao, Sid Kiblawi, Naoto Usuyama +4

Detecting and segmenting small objects, such as lung nodules and tumor lesions, remains a critical challenge in image analysis. These objects often occupy less than 0.1% of an imag…

cs.CV2024

BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once

Theodore Zhao, Yu Gu, Jianwei Yang +12

Biomedical image analysis is fundamental for biomedical discovery in cell biology, pathology, radiology, and many other biomedical domains. Holistic image analysis comprises interd…

cs.CV202414 cited

Foundation Models for Biomedical Image Segmentation: A Survey

Ho Hin Lee, Yu Gu, Theodore Zhao +9

Recent advancements in biomedical image analysis have been significantly driven by the Segment Anything Model (SAM). This transformative technology, originally developed for genera…

cs.CV2023

When an Image is Worth 1,024 x 1,024 Words: A Case Study in Computational Pathology

Wenhui Wang, Shuming Ma, Hanwen Xu +4

This technical report presents LongViT, a vision Transformer that can process gigapixel images in an end-to-end manner. Specifically, we split the gigapixel image into a sequence o…