activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

SigLIP-HD by Fine-to-Coarse Supervision

Lihe Yang, Zhen Zhao, Hengshuang Zhao

High-quality visual representation is a long-standing pursuit in computer vision. In the context of multimodal LLMs (MLLMs), feeding higher-resolution images can produce more fine-…

cs.CV2026

SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry

Zheng Zhang, Lihe Yang, Tianyu Yang +6

We present SCOPE (Scale-Consistent One-Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequences, where existing methods…

cs.CV2025

In Pursuit of Pixel Supervision for Visual Pre-training

Lihe Yang, Shang-Wen Li, Yang Li +5

At the most basic level, pixels are the source of the visual information through which we perceive the world. Pixels contain information at all levels, ranging from low-level attri…

cs.CV2025

Depth Anything with Any Prior

Zehan Wang, Siyu Chen, Lihe Yang +4

This work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in…

cs.CV2025

UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation

Lihe Yang, Zhen Zhao, Hengshuang Zhao

Semi-supervised semantic segmentation (SSS) aims at learning rich visual knowledge from cheap unlabeled images to enhance semantic segmentation capability. Among recent works, UniM…

cs.CV2024

Depth Anything V2

Lihe Yang, Bingyi Kang, Zilong Huang +4

This work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation mo…