4 papers
Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design
Wenxin Wang, Yule Hou, Yu Ji +2
Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even under low-concurrency workloads. We identify…
Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation
Yee Hin Chong, Jiaming Wu, Youhui Zhang +1
Large language models (LLMs) have shown strong empirical gains as self-evolving agents for CUDA kernel generation, driven by feedback-conditioned planning across generations. Howev…
BrainFuse: a unified infrastructure integrating realistic biological modeling and core AI methodology
Baiyu Chen, Yujie Wu, Siyuan Xu +9
Neuroscience and artificial intelligence represent distinct yet complementary pathways to general intelligence. However, amid the ongoing boom in AI research and applications, the…
Pipelining Kruskal's: A Neuromorphic Approach for Minimum Spanning Tree
Yee Hin Chong, Peng Qu, Yuchen Li +1
Neuromorphic computing, characterized by its event-driven computation and massive parallelism, is particularly effective for handling data-intensive tasks in low-power environments…