collaborators

8 papers

cs.DC2026

Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems

Yuchen Fan, Minghong Sun, Jikui Ma +19

AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collec…

cs.IR2026

Uncertainty-Aware Cross-Modal Remote Sensing Image-Text Retrieval via Evidential Learning

Zhuoyue Wang, Xueqian Wang, Gang Li +3

In cross-modal remote sensing image-text retrieval (CMRSITR), test-time remote sensing (RS) images and textual descriptions may deviate from well-curated benchmark conditions due t…

cs.RO2026

The Speedup Paradox: Rethinking Inference Speed-Quality Trade-off in Embodied Tasks

Yujin Wang, Junli Chen, Yixuan Li +4

Embodied foundation models have recently been widely used to improve robot generalization and task success rates. Previous works apply lossy efficient-inference techniques such as…

cs.AR2026

Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators

Zongle Huang, Hongyang Jia, Kaiwei Zou +1

Neural network (NN) accelerators with multi-chip-module (MCM) architectures enable integration of massive computation capability; however, they face challenges of computing resourc…

cs.LG2026

MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE

Zongle Huang, Lei Zhu, Zongyuan Zhan +5

Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional…

cs.LG2025

SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference

Wenxun Wang, Shuchang Zhou, Wenyu Sun +2

Transformers have shown remarkable performance in both natural language processing (NLP) and computer vision (CV) tasks. However, their real-time inference speed and efficiency are…