1 citations · 1 across the 5 of their papers we have counts for
11 papers
Signal or Noise? Understanding Generative Models for Real-World Sensor Time Series
Zitao Shuai, Zongzhe Xu, Yuntian Wu +3
Generative models have changed how machine learning represents complex data distributions, especially in language and vision, yet many real-world systems are observed instead as co…
ELF: Embedded Language Flows
Keya Hu, Linlu Qiu, Yiyang Lu +5
Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing…
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Zhiheng Liu, Weiming Ren, Xiaoke Huang +12
Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the t…
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
Chao Li, Tianhong Li, Sai Vidyaranya Nuthalapati +9
Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes:…
One-step Latent-free Image Generation with Pixel Mean Flows
Yiyang Lu, Susie Lu, Qiao Sun +6
Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operating in a latent space. Recent adva…
Latent Denoising Makes Good Tokenizers
Jiawei Yang, Tianhong Li, Lijie Fan +2
Despite their fundamental role, it remains unclear what properties could make tokenizers more effective for generative modeling. We observe that modern generative models share a co…