14 papers
STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
Ye Wang, Hongjun Wang, Hao Fang +7
Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic con…
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
Yitong Jiang, Hongjun Wang, Collin McCarthy +15
Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining. Subquadratic al…
Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models
Hongjun Wang, Po Hu, Kai Han
Generalized Category Discovery (GCD) aims to categorize unlabelled instances from both known and unknown classes by transferring knowledge from labelled data of known classes. Exis…
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
Inclusion AI, Tiwei Bie, Haoxing Chen +15
We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a natively integrated framework. Its…
MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning
Hongjun Wang, Wei Liu, Weibo Gu +2
Regulating the importance ratio is critical for the training stability of Group Relative Policy Optimization (GRPO) based frameworks. However, prevailing ratio control methods, suc…
Leum-VL Technical Report
Yuxuan He, Chaiming Huang, Yifan Wu +4
A short video succeeds not simply because of what it shows, but because of how it schedules attention -- yet current multimodal models lack the structural grammar to parse or produ…