collaborators

6 papers

cs.CV2026

Tuning-Free Latent Diffusion Models for Ultrahigh-Resolution Image Editing

Wanglong Lu, Lingming Su, Kaijie Shi +4

Recent diffusion-based generative models have shown impressive performance in image generation and editing. However, due to memory limitations and the high cost of collecting high-…

eess.AS2026

SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control

Jeng-Yue Liu, Ting-Chao Hsu, Yen-Tung Yeh +2

Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particular…

cs.SD2026

Timed text extraction from Taiwanese Kua-á-hì TV series

Tzu-Hung Huang, Yun-En Tsai, Yun-Ning Hung +3

Taiwanese opera (Kua-á-hì), a major form of local theatrical tradition, underwent extensive television adaptation notably by pioneers like Iûnn Lē-hua. These videos, while pote…

cs.SD2025

LargeSHS: A large-scale dataset of music adaptation

Chih-Pin Tan, Hsuan-Kai Kao, Li Su +1

Recent advances in AI-based music generation have focused heavily on text-conditioned models, with less attention given to reference-based generation such as song adaptation. To su…

cs.CV2025

TextDoctor: Unified Document Image Inpainting via Patch Pyramid Diffusion Models

Wanglong Lu, Lingming Su, Jingjing Zheng +6

Digital versions of real-world text documents often suffer from issues like environmental corrosion of the original document, low-quality scanning, or human interference. Existing…

cs.SD2024

Distortion Recovery: A Two-Stage Method for Guitar Effect Removal

Ying-Shuo Lee, Yueh-Po Peng, Jui-Te Wu +3

Removing audio effects from electric guitar recordings makes it easier for post-production and sound editing. An audio distortion recovery model not only improves the clarity of th…