3 papers
cs.CV2026
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
Yangming Shi, Shixiang Zhu, Tao Shen +14
We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the m…
cs.CV2025
High-Resolution Image Synthesis via Next-Token Prediction
Dengsheng Chen, Jie Hu, Tiezhu Yue +2
Recently, autoregressive models have demonstrated remarkable performance in class-conditional image generation. However, the application of next-token prediction to high-resolution…
cs.LG2025
Denoising with a Joint-Embedding Predictive Architecture
Dengsheng Chen, Jie Hu, Xiaoming Wei +1
Joint-embedding predictive architectures (JEPAs) have shown substantial promise in self-supervised representation learning, yet their application in generative modeling remains und…