1 citations · 1 across the 2 of their papers we have counts for
5 papers
Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
Ye Qiao, Haocheng Xu, Xiaofan Zhang +1
Extending the context window support of large language models (LLMs) is crucial for tasks with long-distance dependencies. RoPE-based interpolation and extrapolation methods, such…
Characterizing State Space Model and Hybrid Language Model Performance with Long Context
Saptarshi Mitra, Rachid Karami, Haocheng Xu +2
Emerging applications such as AR are driving demands for machine intelligence capable of processing continuous and/or long-context inputs on local devices. However, currently domin…
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
Haocheng Xu, Faraz Tahmasebi, Ye Qiao +3
Recent innovations in Transformer-based large language models have significantly advanced the field of general-purpose neural language understanding and generation. With billions o…
Optimizing High-Level Synthesis Designs with Retrieval-Augmented Large Language Models
Haocheng Xu, Haotian Hu, Sitao Huang
High-level synthesis (HLS) allows hardware designers to create hardware designs with high-level programming languages like C/C++/OpenCL, which greatly improves hardware design prod…
MONAS: Efficient Zero-Shot Neural Architecture Search for MCUs
Ye Qiao, Haocheng Xu, Yifan Zhang +1
Neural Architecture Search (NAS) has proven effective in discovering new Convolutional Neural Network (CNN) architectures, particularly for scenarios with well-defined accuracy opt…