5 papers
LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
Fan Yang, Yuting Su, Xiaobo Wang +7
World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. H…
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
Ce Zheng, Xinghan Wang, Jiahong Ning +3
Federated inference enhances LLM performance in edge computing through weighted averaging of distributed model predictions. However, autoregressive LLM inference requires frequent…
Tool-RoCo: An Agent-as-Tool Self-organization Large Language Model Benchmark in Multi-robot Cooperation
Ke Zhang, Xiaoning Zhao, Ce Zheng +5
This study proposes Tool-RoCo, a novel benchmark for evaluating large language models (LLMs) in long-term multi-agent cooperation based on RoCo, a multi-robot cooperative benchmark…
DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding
Jiahong Ning, Ce Zheng, Tingting Yang
Large language models (LLMs) have transformed natural language processing but face critical deployment challenges in device-edge systems due to resource limitations and communicati…
EdgePrompt: A Distributed Key-Value Inference Framework for LLMs in 6G Networks
Jiahong Ning, Pengyan Zhu, Ce Zheng +3
As sixth-generation (6G) networks advance, large language models (LLMs) are increasingly integrated into 6G infrastructure to enhance network management and intelligence. However,…