collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers

Maoliang Li, Haojing Chen, Jiayu Chen +4

Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bottlenecked by redundant computation acro…

cs.LG2026

DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models

Zihao Zheng, Hangyu Cao, Sicheng Tian +9

Vision-Language-Action (VLA) models are dominant in embodied intelligence but are constrained by inference overheads. While model quantization alleviates these bottlenecks for edge…

cs.LG2026

Vision Transformers that Never Stop Learning

Caihao Sun, Mingqi Yuan, Shiyuan Wang +1

Loss of plasticity refers to the progressive inability of a model to adapt to new tasks and poses a fundamental challenge for continual learning. While this phenomenon has been ext…

cs.LG2026

DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation

Zihao Zheng, Xiuping Cui, Size Zheng +4

As the Mix-of-Experts (MoE) architecture increases the number of parameters in large models, there is an even greater need for model quantization. However, existing quantization me…

cs.LG2025

FedHQ: Hybrid Runtime Quantization for Federated Learning

Zihao Zheng, Ziyao Wang, Xiuping Cui +6

Federated Learning (FL) is a decentralized model training approach that preserves data privacy but struggles with low efficiency. Quantization, a powerful training optimization tec…

cs.LG2025

Threshold Neuron: A Brain-inspired Artificial Neuron for Efficient On-device Inference

Zihao Zheng, Yuanchun Li, Jiayu Chen +3

Enhancing the computational efficiency of on-device Deep Neural Networks (DNNs) remains a significant challengein mobile and edge computing. As we aim to execute increasingly compl…