7 papers
Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
Zhijun Tu, Jian Li, Yuanyuan Xi +5
1-bit LLM quantization offers significant advantages in reducing storage and computational costs. However, existing methods typically train 1-bit LLMs from scratch, failing to full…
EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution
Haizhen Xie, Kunpeng Du, Qiangyu Yan +5
Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. While T2I models have traditionally…
SaD: A Scenario-Aware Discriminator for Speech Enhancement
Xihao Yuan, Siqi Liu, Yan Chen +4
Generative adversarial network-based models have shown remarkable performance in the field of speech enhancement. However, the current optimization strategies for these models pred…
DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution
Miaomiao Cai, Simiao Li, Wei Li +4
Recent advances in diffusion models have improved Real-World Image Super-Resolution (Real-ISR), but existing methods lack human feedback integration, risking misalignment with huma…
Autoregressive Image Generation with Vision Full-view Prompt
Miaomiao Cai, Guanjie Wang, Wei Li +4
In autoregressive (AR) image generation, models based on the 'next-token prediction' paradigm of LLMs have shown comparable performance to diffusion models by reducing inductive bi…
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
Xihao Yuan, Siqi Liu, Hanting Chen +3
Deep learning-based speech enhancement (SE) models have recently outperformed traditional techniques, yet their deployment on resource-constrained devices remains challenging due t…