6 papers
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
Xinming Wei, Jiahao Zhang, Haoran Li +6
Personal LLM agents increasingly combine foreground reactive interactions with background proactive monitoring, forming long-lived, stateful LLM flows that interleave prefill and t…
FedHQ: Hybrid Runtime Quantization for Federated Learning
Zihao Zheng, Ziyao Wang, Xiuping Cui +6
Federated Learning (FL) is a decentralized model training approach that preserves data privacy but struggles with low efficiency. Quantization, a powerful training optimization tec…
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
Feng Jiang, Zihao Zheng, Xiuping Cui +3
With the development of Embodied Artificial intelligence, the end-to-end control policy such as Vision-Language-Action (VLA) model has become the mainstream. Existing VLA models fa…
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
Zihao Zheng, Xiuping Cui, Size Zheng +4
As the Mix-of-Experts (MoE) architecture increases the number of parameters in large models, there is an even greater need for model quantization. However, existing quantization me…
Data and System Perspectives of Sustainable Artificial Intelligence
Tao Xie, David Harel, Dezhi Ran +11
Sustainable AI is a subfield of AI for concerning developing and using AI systems in ways of aiming to reduce environmental impact and achieve sustainability. Sustainable AI is inc…
Threshold Neuron: A Brain-inspired Artificial Neuron for Efficient On-device Inference
Zihao Zheng, Yuanchun Li, Jiayu Chen +3
Enhancing the computational efficiency of on-device Deep Neural Networks (DNNs) remains a significant challengein mobile and edge computing. As we aim to execute increasingly compl…