6 papers
Domain Transfer Becomes Identifiable via a Single Alignment
Sagar Shrestha, Subash Timilsina, Hoang-Son Nguyen +1
Domain transfer (DT) maps source to target distributions and supports tasks such as unsupervised image-to-image translation, single-cell analysis, and cross-platform medical imagin…
Content-Style Identification via Differential Independence
Subash Timilsina, Hoang-Son Nguyen, Sagar Shrestha +1
Generative analysis often models multi-domain observations as nonlinear mixtures of domain-invariant content variables and domain-specific style variables. Identifying both factors…
FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion
Lichen Ma, Zipeng Guo, Yu He +5
To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged as a compelling end-to-end pa…
World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2…
DuoGen: Towards General Purpose Interleaved Multimodal Generation
Min Shi, Xiaohui Zeng, Jiannan Huang +13
Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts f…
Plenoptic Video Generation
Xiao Fu, Shitao Tang, Min Shi +5
Camera-controlled generative video re-rendering methods, such as ReCamMaster, have achieved remarkable progress. However, despite their success in single-view setting, these works…