collaborators

6 papers

cs.AR2026

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference

Zheng Liu, Zeyu Guo, Zihan Liu +9

Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and genera…

cs.AR2026

NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing

Cheng Zou, Shuo Yang, Chen Nie +6

As large language models (LLMs) continue to advance, retrieval-augmented generation (RAG) has become the key mechanism for expanding model knowledge and reducing hallucinations. Ce…

cs.AR2026

ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing

Kang You, Chen Nie, Lee Jun Yan +6

Spiking neural networks (SNNs) exploit event-driven and addition-only computation to substantially improve efficiency for intelligent computation. A key temporal property of SNNs,…

cs.GR2025

SeeLe: A Unified Acceleration Framework for Real-Time Gaussian Splatting

Xiaotong Huang, He Zhu, Zihan Liu +6

3D Gaussian Splatting (3DGS) has become a crucial rendering technique for many real-time applications. However, the limited hardware resources on today's mobile platforms hinder th…

cs.AR2025

Splatonic: Architecture Support for 3D Gaussian Splatting SLAM via Sparse Processing

Xiaotong Huang, He Zhu, Tianrui Ma +8

3D Gaussian splatting (3DGS) has emerged as a promising direction for SLAM due to its high-fidelity reconstruction and rapid convergence. However, 3DGS-SLAM algorithms remain impra…

cs.AR2025

StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic Termination

Yu Feng, Zheng Liu, Weikai Lin +6

Point clouds are increasingly important in intelligent applications, but frequent off-chip memory traffic in accelerators causes pipeline stalls and leads to high energy consumptio…