6 papers
Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8
Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive ope…
EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution
Haizhen Xie, Kunpeng Du, Qiangyu Yan +5
Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. While T2I models have traditionally…
SaD: A Scenario-Aware Discriminator for Speech Enhancement
Xihao Yuan, Siqi Liu, Yan Chen +4
Generative adversarial network-based models have shown remarkable performance in the field of speech enhancement. However, the current optimization strategies for these models pred…
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
Xihao Yuan, Siqi Liu, Hanting Chen +3
Deep learning-based speech enhancement (SE) models have recently outperformed traditional techniques, yet their deployment on resource-constrained devices remains challenging due t…
Instruct-IPT: All-in-One Image Processing Transformer via Weight Modulation
Yuchuan Tian, Jianhong Han, Hanting Chen +5
Due to the unaffordable size and intensive computation costs of low-level vision models, All-in-One models that are designed to address a handful of low-level vision tasks simultan…
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
Yuchuan Tian, Zhijun Tu, Hanting Chen +3
Diffusion Transformers (DiTs) introduce the transformer architecture to diffusion tasks for latent-space image generation. With an isotropic architecture that chains a series of tr…