2 papers
cs.CV2026
RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
Guanfang Dong, Luke Schultz, Negar Hassanpour +1
Semantic-rich features from Vision Foundation Models (VFMs) have been leveraged to enhance Latent Diffusion Models (LDMs). However, raw VFM features are typically high-dimensional…
cs.CV2024
Accelerating Inference of Networks in the Frequency Domain
Chenqiu Zhao, Guanfang Dong, Anup Basu
It has been demonstrated that networks' parameters can be significantly reduced in the frequency domain with a very small decrease in accuracy. However, given the cost of frequency…