activity
20242026
collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs

Ye Wang, Hongjun Wang, Hao Fang +7

Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic con…

cs.CV2026

Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders

Yitong Jiang, Hongjun Wang, Collin McCarthy +15

Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining. Subquadratic al…

cs.CV2026

Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models

Hongjun Wang, Po Hu, Kai Han

Generalized Category Discovery (GCD) aims to categorize unlabelled instances from both known and unknown classes by transferring knowledge from labelled data of known classes. Exis…

cs.CV2026

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model

Inclusion AI, Tiwei Bie, Haoxing Chen +15

We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a natively integrated framework. Its…

cs.CV2025

Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation

Weining Ren, Hongjun Wang, Xiao Tan +1

We present Fin3R, a simple, effective, and general fine-tuning method for feed-forward 3D reconstruction models. The family of feed-forward reconstruction model regresses pointmap…

cs.CV2025

Panoptic Captioning: An Equivalence Bridge for Image and Text

Kun-Yu Lin, Hongjun Wang, Weining Ren +1

This work introduces panoptic captioning, a novel task striving to seek the minimum text equivalent of images, which has broad potential applications. We take the first step toward…