most citedMONAS: Efficient Zero-Shot Neural Architecture Search for MCUs

1 citations · 2 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG2025

Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs

Ye Qiao, Sitao Huang

Extending LLM context windows is crucial for long range tasks. RoPE-based position interpolation (PI) methods like linear and frequency-aware scaling extend input lengths without r…

cs.AR2025

COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference

Ye Qiao, Zhiheng Chen, Yian Wang +3

Transformer-based models have demonstrated superior performance in various fields, including natural language processing and computer vision. However, their enormous model size and…

cs.AR20251 cited

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs

Ye Qiao, Zhiheng Chen, Yifan Zhang +2

Deploying large language models (LLMs) on edge platforms is challenged by their high computational and memory demands. Although recent low-bit quantization methods (e.g., BitNet, D…

cs.LG20241 cited

MONAS: Efficient Zero-Shot Neural Architecture Search for MCUs

Ye Qiao, Haocheng Xu, Yifan Zhang +1

Neural Architecture Search (NAS) has proven effective in discovering new Convolutional Neural Network (CNN) architectures, particularly for scenarios with well-defined accuracy opt…

cs.LG2024

MicroNAS: Zero-Shot Neural Architecture Search for MCUs

Ye Qiao, Haocheng Xu, Yifan Zhang +1

Neural Architecture Search (NAS) effectively discovers new Convolutional Neural Network (CNN) architectures, particularly for accuracy optimization. However, prior approaches often…