collaborators

7 papers

cs.RO2026

RhinoVLA Technical Report

Huixi Technology, :, Chen Zhang +13

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but real-time deployment on edge hardware remains challenging. In this work, we identify V…

cs.GR2026

Efficient 3D Gaussian Splatting with Axis-Shared Rasterization and Order-independent Transmittance

Zhican Wang, Guanghui He, Lingjun Gao +6

3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, combining high-quality reconstruction with efficient rendering. It has been widely adopte…

cs.AR2026

MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems

Zhuoshan Zhou, Chen Zhang, Shuyi Zhang +10

The Mixture-of-Experts (MoE) architecture is crucial for scaling large language models, but its scalability is severely limited by inter-GPU communication bottlenecks in multi-GPU…

cs.AR2026

Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs

Qijun Zhang, Chen Zhang, Zhuoshan Zhou +10

Mixture-of-Experts (MoE) has been adopted by many leading large models to reduce computational requirements. However, frequent inter-GPU communication in MoE expert parallelism (EP…

cs.AR2026

DS2SC-Agent: A Multi-Agent Automated Pipeline for Rapid Chiplet Model Generation

Yiwei Wu, Yifan Wu, Yunhao Xiong +6

Constructing behavioral-level chiplet models (e.g., SystemC) is crucial for early-stage heterogeneous architecture exploration. Traditional manual modeling is notoriously time-cons…

cs.AR2025

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations

Zhican Wang, Guanghui He, Hongxiang Fan

The emergence of diffusion models has significantly advanced generative AI, improving the quality, realism, and creativity of image and video generation. Among them, Stable Diffusi…