8 papers
Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems
Yuchen Fan, Minghong Sun, Jikui Ma +19
AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collec…
Uncertainty-Aware Cross-Modal Remote Sensing Image-Text Retrieval via Evidential Learning
Zhuoyue Wang, Xueqian Wang, Gang Li +3
In cross-modal remote sensing image-text retrieval (CMRSITR), test-time remote sensing (RS) images and textual descriptions may deviate from well-curated benchmark conditions due t…
The Speedup Paradox: Rethinking Inference Speed-Quality Trade-off in Embodied Tasks
Yujin Wang, Junli Chen, Yixuan Li +4
Embodied foundation models have recently been widely used to improve robot generalization and task success rates. Previous works apply lossy efficient-inference techniques such as…
Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators
Zongle Huang, Hongyang Jia, Kaiwei Zou +1
Neural network (NN) accelerators with multi-chip-module (MCM) architectures enable integration of massive computation capability; however, they face challenges of computing resourc…
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
Zongle Huang, Lei Zhu, Zongyuan Zhan +5
Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional…
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
Wenxun Wang, Shuchang Zhou, Wenyu Sun +2
Transformers have shown remarkable performance in both natural language processing (NLP) and computer vision (CV) tasks. However, their real-time inference speed and efficiency are…