2 papers
cs.AR2026
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
Tao Zhang, Rui Ma, Shuotao Xu +2
GPU design space exploration (DSE) for modern AI workloads, such as Large-Language Model (LLM) inference, is challenging because of GPUs' vast, multi-modal design spaces, high simu…
cs.DC2025
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
Jiamin Li, Lei Qu, Tao Zhang +4
This document presents a vision for a novel AI infrastructure design that has been initially validated through inference simulations on state-of-the-art large language models. Adva…