activity
20242026
collaborators

8 papers

cs.CV2026

SigLIP-HD by Fine-to-Coarse Supervision

Lihe Yang, Zhen Zhao, Hengshuang Zhao

High-quality visual representation is a long-standing pursuit in computer vision. In the context of multimodal LLMs (MLLMs), feeding higher-resolution images can produce more fine-…

cs.CV2025

UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation

Lihe Yang, Zhen Zhao, Hengshuang Zhao

Semi-supervised semantic segmentation (SSS) aims at learning rich visual knowledge from cheap unlabeled images to enhance semantic segmentation capability. Among recent works, UniM…

cs.CV2024

SOEDiff: Efficient Distillation for Small Object Editing

Yiming Wu, Qihe Pan, Zhen Zhao +3

In this paper, we delve into a new task known as small object editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remark…

cs.CV2024

DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model

Zhenhua Xu, Yujia Zhang, Enze Xie +5

Multimodal large language models (MLLMs) have emerged as a prominent area of interest within the research community, given their proficiency in handling and reasoning with non-text…

cs.CV2024

GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation

Zhanyu Wang, Longyue Wang, Zhen Zhao +7

While the recent advances in Multimodal Large Language Models (MLLMs) constitute a significant leap forward in the field, these models are predominantly confined to the realm of in…

cs.CV2024

Depth Anything V2

Lihe Yang, Bingyi Kang, Zilong Huang +4

This work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation mo…