6 papers
Tuning-Free Latent Diffusion Models for Ultrahigh-Resolution Image Editing
Wanglong Lu, Lingming Su, Kaijie Shi +4
Recent diffusion-based generative models have shown impressive performance in image generation and editing. However, due to memory limitations and the high cost of collecting high-…
SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control
Jeng-Yue Liu, Ting-Chao Hsu, Yen-Tung Yeh +2
Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particular…
Timed text extraction from Taiwanese Kua-á-hì TV series
Tzu-Hung Huang, Yun-En Tsai, Yun-Ning Hung +3
Taiwanese opera (Kua-á-hì), a major form of local theatrical tradition, underwent extensive television adaptation notably by pioneers like Iûnn LÄ-hua. These videos, while pote…
LargeSHS: A large-scale dataset of music adaptation
Chih-Pin Tan, Hsuan-Kai Kao, Li Su +1
Recent advances in AI-based music generation have focused heavily on text-conditioned models, with less attention given to reference-based generation such as song adaptation. To su…
TextDoctor: Unified Document Image Inpainting via Patch Pyramid Diffusion Models
Wanglong Lu, Lingming Su, Jingjing Zheng +6
Digital versions of real-world text documents often suffer from issues like environmental corrosion of the original document, low-quality scanning, or human interference. Existing…
Distortion Recovery: A Two-Stage Method for Guitar Effect Removal
Ying-Shuo Lee, Yueh-Po Peng, Jui-Te Wu +3
Removing audio effects from electric guitar recordings makes it easier for post-production and sound editing. An audio distortion recovery model not only improves the clarity of th…