3 papers
cs.AI2026
FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference
Qihu Xie, Ziwei Li, Yi Kang
Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memo…
cs.CV2025
HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
Shizhuo Mao, Hongtao Zou, Qihu Xie +2
Diffusion models have demonstrated significant applications in the field of image generation. However, their high computational and memory costs pose challenges for deployment. Mod…
cs.LG2025
SBS: Enhancing Parameter-Efficiency of Neural Representations for Neural Networks via Spectral Bias Suppression
Qihu Xie, Yuan Li, Yi Kang
Implicit neural representations have recently been extended to represent convolutional neural network weights via neural representation for neural networks, offering promising para…