1 citations · 1 across the 22 of their papers we have counts for
5 papers · 1 filter
DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU
Weizhe Chen, Miao Zhang, Junpeng Jiang +3
Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality, making hybrid architecture des…
Boost Post-Training Quantization via Null Space Optimization for Large Language Models
Jiaqi Zhao, Miao Zhang, Deng Xiang +3
Existing post-training quantization methods for large language models (LLMs) offer remarkable success. However, the increasingly marginal performance gains suggest that existing qu…
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
Jiaqi Zhao, Miao Zhang, Ming Wang +5
Large Language Models (LLMs) suffer severe performance degradation when facing extremely low-bit (sub 2-bit) quantization. Several existing sub 2-bit post-training quantization (PT…
SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting
Haomiao Qiu, Miao Zhang, Ziyue Qiao +3
Continual Learning requires a model to learn multiple tasks in sequence while maintaining both stability:preserving knowledge from previously learned tasks, and plasticity:effectiv…
Content-aware Balanced Spectrum Encoding in Masked Modeling for Time Series Classification
Yudong Han, Haocong Wang, Yupeng Hu +3
Due to the superior ability of global dependency, transformer and its variants have become the primary choice in Masked Time-series Modeling (MTM) towards time-series classificatio…