1 citations · 2 across the 5 of their papers we have counts for
5 papers
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
Ye Qiao, Sitao Huang
Extending LLM context windows is crucial for long range tasks. RoPE-based position interpolation (PI) methods like linear and frequency-aware scaling extend input lengths without r…
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
Ye Qiao, Zhiheng Chen, Yian Wang +3
Transformer-based models have demonstrated superior performance in various fields, including natural language processing and computer vision. However, their enormous model size and…
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
Ye Qiao, Zhiheng Chen, Yifan Zhang +2
Deploying large language models (LLMs) on edge platforms is challenged by their high computational and memory demands. Although recent low-bit quantization methods (e.g., BitNet, D…
MONAS: Efficient Zero-Shot Neural Architecture Search for MCUs
Ye Qiao, Haocheng Xu, Yifan Zhang +1
Neural Architecture Search (NAS) has proven effective in discovering new Convolutional Neural Network (CNN) architectures, particularly for scenarios with well-defined accuracy opt…
MicroNAS: Zero-Shot Neural Architecture Search for MCUs
Ye Qiao, Haocheng Xu, Yifan Zhang +1
Neural Architecture Search (NAS) effectively discovers new Convolutional Neural Network (CNN) architectures, particularly for accuracy optimization. However, prior approaches often…