2 papers
cs.DC2026
Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap
Ziyu Huang, Yangjie Zhou, Chenhao Zhu +14
Mixture-of-Experts (MoE) inference under expert parallelism (EP) turns each MoE layer into a distributed computation with costly dispatch and combine communication. State-of-the-ar…
cs.CL2026
Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
Qi Chen, Yunfei Chu, Haolin He +15
Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed tex…