1 citations · 1 across the 9 of their papers we have counts for
9 papers · 1 filter
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
Penghao Wu, Haiwen Diao, Weichen Fan +3
While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, functional unification does not guarantee learning synergy: the tw…
Apple-: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Runmao Yao, Kairui Hu, Yukang Cao +11
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausi…
Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion
Weichen Fan, Haiwen Diao, Penghao Wu +1
Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoiser is strongly frequency dependent. Under rectified-flow diffu…
From Pixels to Words -- Towards Native One-Vision Models at Scale
Haiwen Diao, Jiahao Wang, Penghao Wu +18
Current vision-language models (VLMs) typically stitch together separate image encoders and language decoders via multi-stage alignment, a modular framework that inevitably fragmen…
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Haiwen Diao, Penghao Wu, Hanming Deng +55
Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fra…
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
Weichen Fan, Haiwen Diao, Quan Wang +2
Deep representations across modalities are inherently intertwined. In this paper, we systematically analyze the spectral characteristics of various semantic and pixel encoders. Int…