2 papers
cs.AI2026
SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning
Yujie Zhang, Bin Gao, Tulika Mitra
Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usa…
cs.AR2026
A Data-Driven Dynamic Execution Orchestration Architecture
Zhenyu Bai, Pranav Dangi, Rohan Juneja +4
Domain-specific accelerators deliver exceptional performance on their target workloads through fabrication-time orchestrated datapaths. However, such specialized architectures ofte…