10 papers
Stable FP4 Training via Transposition-Invariant Block Quantization
Mehdi Rahimifar, Amin Darabi, Mehran Taghian Jazi +6
Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging…
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
Nuoyan Zhou, Zhijun Tu, Lei Yu +4
Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong capabilities in image generation. However,…
Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
Zhijun Tu, Jian Li, Yuanyuan Xi +5
1-bit LLM quantization offers significant advantages in reducing storage and computational costs. However, existing methods typically train 1-bit LLMs from scratch, failing to full…
ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices
Kunpeng Du, Haizhen Xie, Sen Lu +11
The Diffusion Transformer (DiT) architecture is the state-of-the-art paradigm for high-fidelity image generation, underpinning models like Stable Diffusion-3 and FLUX.1. However, d…
Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution
Xiao He, Zhijun Tu, Kun Cheng +4
The demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their a…
One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation
Xue Wu, Jingwei Xin, Zhijun Tu +4
Diffusion-based models have been widely used in various visual generation tasks, showing promising results in image super-resolution (SR), while typically being limited by dozens o…