84 citations · 147 across the 28 of their papers we have counts for
6 papers · 1 filter
Instruct-IPT: All-in-One Image Processing Transformer via Weight Modulation
Yuchuan Tian, Jianhong Han, Hanting Chen +5
Due to the unaffordable size and intensive computation costs of low-level vision models, All-in-One models that are designed to address a handful of low-level vision tasks simultan…
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
Yuchuan Tian, Zhijun Tu, Hanting Chen +3
Diffusion Transformers (DiTs) introduce the transformer architecture to diffusion tasks for latent-space image generation. With an isotropic architecture that chains a series of tr…
Distilling Semantic Priors from SAM to Efficient Image Restoration Models
Quan Zhang, Xiaoyu Liu, Wei Li +6
In image restoration (IR), leveraging semantic priors from segmentation models has been a common approach to improve performance. The recent segment anything model (SAM) has emerge…
DiJiang: Efficient Large Language Models through Compact Kernelization
Hanting Chen, Zhicheng Liu, Xutao Wang +2
In an effort to reduce the computational load of Transformers, research on linear attention has gained significant momentum. However, the improvement strategies for attention mecha…
IPT-V2: Efficient Image Processing Transformer using Hierarchical Attentions
Zhijun Tu, Kunpeng Du, Hanting Chen +4
Recent advances have demonstrated the powerful capability of transformer architecture in image restoration. However, our analysis indicates that existing transformerbased methods c…
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Jianyuan Guo, Hanting Chen, Chengcheng Wang +3
Recent advancements in large language models have sparked interest in their extraordinary and near-superhuman capabilities, leading researchers to explore methods for evaluating an…