5 papers
What Is Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution
Xingsong Ye, Yongkun Du, JiaXin Zhang +3
Large-scale and categorical-balanced text data is essential for training effective Scene Text Recognition (STR) models, which is hard to achieve when collecting real data. Syntheti…
NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
Tianlin Pan, Jiayi Dai, Chenpu Yuan +7
Recent video editing models have achieved impressive results, but most still require large-scale paired datasets. Collecting such naturally aligned pairs at scale remains highly ch…
Improving Reconstruction of Representation Autoencoder
Siyu Liu, Chujie Qin, Hubery Yin +6
Recent work leverages Vision Foundation Models as image encoders to boost the generative performance of latent diffusion models (LDMs), as their semantic feature distributions are…
Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
Haotian Dong, Wenjing Wang, Chen Li +2
Generating RGB-A videos, which include alpha channels for transparency, has wide applications. However, current methods often suffer from low quality due to confusion between RGB a…
VACoT: Rethinking Visual Data Augmentation with VLMs
Zhengzhuo Xu, Chong Sun, SiNan Du +3
While visual data augmentation remains a cornerstone for training robust vision models, it has received limited attention in visual language models (VLMs), which predominantly rely…