activity
20242026
collaborators

9 papers

cs.LG2026

Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding

Zhongyu Xiao, Zhiwei Hao, Jianyuan Guo +4

Diffusion Large Language Models (dLLMs) offer a compelling paradigm for natural language generation, leveraging parallel decoding and bidirectional attention to achieve superior gl…

cs.AI2026

From Question Answering to Task Completion: A Survey on Agent System and Harness Design

Jianyuan Guo, Zhiwei Hao, Chengcheng Wang +14

LLM-based agents mark a shift from passive question answering to active task completion: they perceive environments, invoke tools, maintain state, and act over extended horizons. A…

cs.LG2026

LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing

Jiawei Hao, Zhiwei Hao, Jianyuan Guo +4

Mixture-of-Experts (MoE) based Large Language Models (LLMs) have demonstrated impressive performance and computational efficiency. However, their deployment is often constrained by…

cs.CV2025

ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters

Zhiwei Hao, Jianyuan Guo, Li Shen +4

Recent advancements in vision transformers (ViTs) have demonstrated that larger models often achieve superior performance. However, training these models remains computationally in…

cs.DC2025

CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference

Guanyu Xu, Zhiwei Hao, Li Shen +5

The impressive performance of transformer models has sparked the deployment of intelligent applications on resource-constrained edge devices. However, ensuring high-quality service…

cs.LG2025

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities

Zhiwei Hao, Jianyuan Guo, Li Shen +6

Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources required for their training present a signific…