3 papers
cs.SD2026
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Haowei Lou, Hye-Young Paik, Dai Jia +2
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for spe…
cs.LG2026
SAD-LoRA: Spectral Alignment for Low-Rank Knowledge Distillation
Omer Tariq, Syed Muhammad Raza, Jeongbae Son
Distilling a fine-tuned teacher into a LoRA-adapted student is a standard recipe for parameter-efficient compression, but output-level KD does not explicitly control which rank-…
cs.CV2026
Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
Omer Tariq, Syed Muhammad Raza, Jeongbae Son
Video summarization aims to produce a compact representation of a long video by selecting a subset of temporally important segments that best reflect human preferences. This task i…