5 papers
TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration
Linye Wei, Zixiang Luo, Pingzhi Tang +1
Diffusion large language models (dLLMs) have recently gained significant attention due to their inherent support for parallel decoding. Building on this paradigm, Mixture-of-Expert…
Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
Linye Wei, Wenjue Chen, Pingzhi Tang +4
Diffusion-based large language models (dLLMs) have recently gained significant attention for their exceptional performance and inherent potential for parallel decoding. Existing fr…
No Redundancy, No Stall: Lightweight Streaming 3D Gaussian Splatting for Real-time Rendering
Linye Wei, Jiajun Tang, Fan Fei +3
3D Gaussian Splatting (3DGS) enables high-quality rendering of 3D scenes and is getting increasing adoption in domains like autonomous driving and embodied intelligence. However, 3…
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
Linye Wei, Shuzhang Zhong, Songqiang Xu +3
Large language model (LLM)-based automatic speech recognition (ASR) has recently attracted a lot of attention due to its high recognition accuracy and enhanced multi-dialect suppor…
VR-YOLO: Enhancing PCB Defect Detection with Viewpoint Robustness Based on YOLO
Hengyi Zhu, Linye Wei, He Li
The integration of large-scale circuits and systems emphasizes the importance of automated defect detection of electronic components. The YOLO image detection model has been used t…