4 papers
DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference
Xiaoyang Lu, Belthangady Akash Vi Narayana Pai, Xian-He Sun
Mixture-of-Experts (MoE) models enable efficient scaling of large language model (LLM) inference but suffer from substantial data-movement overhead when deployed on neural processi…
VIPER: Architecture-Aware Performance Modeling for Processing-in-Memory Design-Space Exploration
Haoran Geng, Tomas Sousa Pereira, Xiaoyang Lu +3
Processing-in-Memory (PIM) promises to reduce data movement overhead by executing computation in or near memory, but its realized application speedup remains highly design-dependen…
SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale
Jiaqi Liu, Maolin Ran, Xiaoyang Lu +5
Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity has become a critical bottlenec…
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
Maolin Ran, Xiaoyang Lu, Jiaqi Liu +5
Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial…