3 papers
eess.IV2026
TaQ-DiT: Time-aware Quantization for Diffusion Transformers
Xinyan Liu, Huihong Shi, Yang Xu +1
Transformer-based diffusion models, dubbed Diffusion Transformers (DiTs), have achieved state-of-the-art performance in image and video generation tasks. However, their large model…
eess.AS2026
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
Leyan Yang, Ronghui Hu, Yang Xu +1
Recent advancements in end-to-end neural speech codecs enable compressing audio at extremely low bitrates while maintaining high-fidelity reconstruction. Meanwhile, low computation…
cs.CV2025
Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained Experts
Yangyang Xu, Xi Ye, Duo Su
Multi-task learning (MTL) for dense prediction has shown promising results but still faces challenges in balancing shared representations with task-specific specialization. In this…