most citedScale-Aware Modulation Meet Transformer

11 citations · 16 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2023

BeautifulPrompt: Towards Automatic Prompt Engineering for Text-to-Image Synthesis

Tingfeng Cao, Chengyu Wang, Bingyan Liu +3

Recently, diffusion-based deep generative models (e.g., Stable Diffusion) have shown impressive results in text-to-image synthesis. However, current text-to-image models often requ…

cs.CV20232 cited

EasyPhoto: Your Smart AI Photo Generator

Ziheng Wu, Jiaqi Xu, Xinyi Zou +3

Stable Diffusion web UI (SD-WebUI) is a comprehensive project that provides a browser interface based on Gradio library for Stable Diffusion models. In this paper, We propose a nov…

cs.CV20231 cited

DualToken-ViT: Position-aware Efficient Vision Transformer with Dual Token Fusion

Zhenzhen Chu, Jiayu Chen, Cen Chen +4

Self-attention-based vision transformers (ViTs) have emerged as a highly competitive architecture in computer vision. Unlike convolutional neural networks (CNNs), ViTs are capable…

cs.CV2023

DiffSynth: Latent In-Iteration Deflickering for Realistic Video Synthesis

Zhongjie Duan, Lizhou You, Chengyu Wang +4

In recent years, diffusion models have emerged as the most powerful approach in image synthesis. However, applying these models directly to video synthesis presents challenges, as…

cs.CV202311 cited

Scale-Aware Modulation Meet Transformer

Weifeng Lin, Ziheng Wu, Jiayu Chen +2

This paper presents a new vision Transformer, Scale-Aware Modulation Transformer (SMT), that can handle various downstream tasks efficiently by combining the convolutional network…

cs.CV20232 cited

SC-ML: Self-supervised Counterfactual Metric Learning for Debiased Visual Question Answering

Xinyao Shu, Shiyang Yan, Xu Yang +3

Visual question answering (VQA) is a critical multimodal task in which an agent must answer questions according to the visual cue. Unfortunately, language bias is a common problem…