5 papers
DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization
Minnan Pei, Gang Li, Zeyu Zhu +7
3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: incre…
A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference
Zhuoran Song, Haozhe Jiang, Chunyu Qi +4
Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators,…
LATTICE: Constraint-Directed Scheduling, Memory Planning, and Pipeline Refinement for NPUs
Runhao Liu, Minman Pei, Peng Zheng +7
General-purpose NPUs execute fine-grained command DAGs across heterogeneous compute and memory-transfer engines backed by finite, explicitly managed on-chip memories. This executio…
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
Zeyu Zhu, Gang Li, Peisong Wang +5
Mixture of Experts (MoE) architectures significantly enhance the capacity of LLMs without proportional increases in computation, but at the cost of a vast parameter size. Offloadin…
GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing
Minnan Pei, Gang Li, Junwen Si +6
3D Gaussian Splatting (3DGS) has emerged as a leading neural rendering technique for high-fidelity view synthesis, prompting the development of dedicated 3DGS accelerators for reso…