activity
20202025
most citedFocal Self-attention for Local-Global Interactions in Vision Transformers

268 citations · 470 across the 20 of their papers we have counts for

collaborators
Showing cs.CVShow all

27 papers · 1 filter

cs.CV2025

Simplifying DINO via Coding Rate Regularization

Ziyang Wu, Jingyuan Zhang, Druv Pai +5

DINO and DINOv2 are two model families being widely used to learn representations from unlabeled imagery data at large scales. Their learned representations often enable state-of-t…

cs.CV2024

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

Jiuhai Chen, Jianwei Yang, Haiping Wu +4

We present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2, a generative vision foundation model.…

cs.CV2024

DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Lingchen Meng, Jianwei Yang, Rui Tian +4

Most large multimodal models (LMMs) are implemented by feeding visual tokens as a sequence into the first layer of a large language model (LLM). The resulting architecture is simpl…

cs.CV2024

BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once

Theodore Zhao, Yu Gu, Jianwei Yang +12

Biomedical image analysis is fundamental for biomedical discovery in cell biology, pathology, radiology, and many other biomedical domains. Holistic image analysis comprises interd…

cs.CV2024

Matryoshka Multimodal Models

Mu Cai, Jianwei Yang, Jianfeng Gao +1

Large Multimodal Models (LMMs) such as LLaVA have shown strong performance in visual-linguistic reasoning. These models first embed images into a fixed large number of visual token…

cs.CV2024

Pix2Gif: Motion-Guided Diffusion for GIF Generation

Hitesh Kandala, Jianfeng Gao, Jianwei Yang

We present Pix2Gif, a motion-guided diffusion model for image-to-GIF (video) generation. We tackle this problem differently by formulating the task as an image translation problem…