2 papers
cs.LG2025
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
Liang Qiao, Yue Dai, Yeqi Huang +3
Multi-Modal Diffusion Transformers (DiTs) demonstrate exceptional capabilities in visual synthesis, yet their deployment remains constrained by substantial computational demands. T…
cs.LG2025
Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program Tuning
Liang Qiao, Jun Shi, Xiaoyu Hao +10
Tensor program tuning is essential for the efficient deployment of deep neural networks. Search-based approaches have demonstrated scalability and effectiveness in automatically fi…