2 papers
cs.DC2026
Beyond Fast Contractions: Attenuation and Recovery of Matrix-Engine Speedups in High-Order Finite Elements
Yinuo Wang, Lin Gan, Tianqi Mao +8
Modern processors increasingly provide matrix engines whose peak arithmetic throughput greatly exceeds conventional SIMD, but scientific applications rarely realize this advantage…
cs.CV2026
DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse
Yuyang Chen, Runxin Zhong, Zan Zong +3
Recent advances in AI-generated content have driven widespread adoption of Diffusion Transformers (DiTs) for high-resolution, long-duration content generation. While parallelizatio…