7 papers
RhinoVLA Technical Report
Huixi Technology, :, Chen Zhang +13
Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but real-time deployment on edge hardware remains challenging. In this work, we identify V…
Efficient 3D Gaussian Splatting with Axis-Shared Rasterization and Order-independent Transmittance
Zhican Wang, Guanghui He, Lingjun Gao +6
3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, combining high-quality reconstruction with efficient rendering. It has been widely adopte…
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
Zhuoshan Zhou, Chen Zhang, Shuyi Zhang +10
The Mixture-of-Experts (MoE) architecture is crucial for scaling large language models, but its scalability is severely limited by inter-GPU communication bottlenecks in multi-GPU…
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
Qijun Zhang, Chen Zhang, Zhuoshan Zhou +10
Mixture-of-Experts (MoE) has been adopted by many leading large models to reduce computational requirements. However, frequent inter-GPU communication in MoE expert parallelism (EP…
DS2SC-Agent: A Multi-Agent Automated Pipeline for Rapid Chiplet Model Generation
Yiwei Wu, Yifan Wu, Yunhao Xiong +6
Constructing behavioral-level chiplet models (e.g., SystemC) is crucial for early-stage heterogeneous architecture exploration. Traditional manual modeling is notoriously time-cons…
SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations
Zhican Wang, Guanghui He, Hongxiang Fan
The emergence of diffusion models has significantly advanced generative AI, improving the quality, realism, and creativity of image and video generation. Among them, Stable Diffusi…