Showing 2026Show all
3 papers · 1 filter
cs.DC2026
Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design
Wenxin Wang, Yule Hou, Yu Ji +2
Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even under low-concurrency workloads. We identify…
cs.AI2026
Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation
Yee Hin Chong, Jiaming Wu, Youhui Zhang +1
Large language models (LLMs) have shown strong empirical gains as self-evolving agents for CUDA kernel generation, driven by feedback-conditioned planning across generations. Howev…
cs.NE2026
BrainFuse: a unified infrastructure integrating realistic biological modeling and core AI methodology
Baiyu Chen, Yujie Wu, Siyuan Xu +9
Neuroscience and artificial intelligence represent distinct yet complementary pathways to general intelligence. However, amid the ongoing boom in AI research and applications, the…