4 papers
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Haowei Lou, Hye-Young Paik, Dai Jia +2
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for spe…
SAD-LoRA: Spectral Alignment for Low-Rank Knowledge Distillation
Omer Tariq, Syed Muhammad Raza, Jeongbae Son
Distilling a fine-tuned teacher into a LoRA-adapted student is a standard recipe for parameter-efficient compression, but output-level KD does not explicitly control which rank-…
Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
Omer Tariq, Syed Muhammad Raza, Jeongbae Son
Video summarization aims to produce a compact representation of a long video by selecting a subset of temporally important segments that best reflect human preferences. This task i…
EdgeDAM: Real-time Object Tracking for Mobile Devices
Syed Muhammad Raza, Syed Murtaza Hussain Abidi, Khawar Islam +2
Single-object tracking (SOT) on edge devices is a critical computer vision task, requiring accurate and continuous target localization across video frames under occlusion, distract…