6 papers · 1 filter
MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers
Maoliang Li, Haojing Chen, Jiayu Chen +4
Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bottlenecked by redundant computation acro…
Beyond Spatial Compression: Interface-Centric Generative States for Open-World 3D Structure
Xiang Chen, Alexander Binder
Current 3D tokenizers largely treat representation as spatial compression: compact codes reconstruct surface geometry, but leave component ownership and attachment validity implici…
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
Zihao Zheng, Hangyu Cao, Sicheng Tian +9
Vision-Language-Action (VLA) models are dominant in embodied intelligence but are constrained by inference overheads. While model quantization alleviates these bottlenecks for edge…
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
Zihao Zheng, Xiuping Cui, Size Zheng +4
As the Mix-of-Experts (MoE) architecture increases the number of parameters in large models, there is an even greater need for model quantization. However, existing quantization me…
FedHQ: Hybrid Runtime Quantization for Federated Learning
Zihao Zheng, Ziyao Wang, Xiuping Cui +6
Federated Learning (FL) is a decentralized model training approach that preserves data privacy but struggles with low efficiency. Quantization, a powerful training optimization tec…
Threshold Neuron: A Brain-inspired Artificial Neuron for Efficient On-device Inference
Zihao Zheng, Yuanchun Li, Jiayu Chen +3
Enhancing the computational efficiency of on-device Deep Neural Networks (DNNs) remains a significant challengein mobile and edge computing. As we aim to execute increasingly compl…