10 papers
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Haowei Lou, Hye-Young Paik, Dai Jia +2
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for spe…
SAD-LoRA: Spectral Alignment for Low-Rank Knowledge Distillation
Omer Tariq, Syed Muhammad Raza, Jeongbae Son
Distilling a fine-tuned teacher into a LoRA-adapted student is a standard recipe for parameter-efficient compression, but output-level KD does not explicitly control which rank-…
NOS-Gate: Queue-Aware Streaming IDS for Consumer Gateways under Timing-Controlled Evasion
Muhammad Bilal, Omer Tariq, Hasan Ahmed
Timing and burst patterns can leak through encryption, and an adaptive adversary can exploit them. This undermines metadata-only detection in a stand-alone consumer gateway. Theref…
Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
Omer Tariq, Syed Muhammad Raza, Jeongbae Son
Video summarization aims to produce a compact representation of a long video by selecting a subset of temporally important segments that best reflect human preferences. This task i…
BARFI-Q: Quantum-Enhanced Block Attention Residual Fusion Framework for Multivariate Time-Series Forecasting in Atom Interferometry
Muhammad Bilal Akram Dastagir, Omer Tariq, Safaa Alqrinawi +3
Atom interferometry generates heterogeneous multivariate temporal streams governed by phase evolution, fringe dynamics, control variables, and auxiliary sensing measurements. Accur…
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
Md Selim Sarowar, Omer Tariq, Sungho Kim
VLA models encode visual observations as 2D patch tokens with no intrinsic geometric structure. We introduce GST-VLA with two contributions. First, the Gaussian Spatial Tokenizer (…