4 papers
Generation Quality-Latency Tradeoff-Aware Inference Offloading for Multimodal LLMs in Cloud-Edge Continuum
Zhongxiao Wang, Yueshen Xu, Xinkui Zhao +2
Beyond pure cloud, some efforts are being made to deploy Large Language Models (LLMs) in edge to accelerate inference response. So the deployment of LLMs in cloud-edge continuum be…
ATOM: Instantiating Budget-Controllable Multi-Agent Collaboration via Nucleus-Electron Hierarchy
Xinkui Zhao, Sai Liu, Yifan Zhang +6
Large Language Model (LLM)-based multi-agent systems rely on optimized collaboration topologies to balance performance and communication costs. However, current methods struggle wi…
StackPilot: Autonomous Function Agents for Scalable and Environment-Free Code Execution
Xinkui Zhao, Yifan Zhang, Zhengyi Zhou +1
Recent advances in large language models (LLMs) have substantially enhanced automated code generation across a wide range of programming languages. Nonetheless, verifying the corre…
TRAIL: Joint Inference and Refinement of Knowledge Graphs with Large Language Models
Xinkui Zhao, Haode Li, Yifan Zhang +2
Recent advances in large language models (LLMs) have unlocked powerful reasoning and decision-making capabilities. However, their inherent dependence on static parametric memory fu…