5 papers
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
Tingyang Sun, Ting He, Bo Ji +1
Large language models have demonstrated extraordinary performance in many AI tasks but are expensive to use, even after training, due to their requirement of high-end GPUs. Recentl…
iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM Inference
Wei Fan, JinYi Yoon, Bo Ji
Large Language Model (LLM) agent systems have advanced rapidly, driven by their strong generalization in zero-shot settings. To further enhance reasoning and accuracy on complex ta…
S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
JinYi Yoon, JiHo Lee, Ting He +2
With the advancement of Artificial Intelligence (AI) towards multiple modalities (language, vision, speech, etc.), multi-modal models have increasingly been used across various app…
P3SL: Personalized Privacy-Preserving Split Learning on Heterogeneous Edge Devices
Wei Fan, JinYi Yoon, Xiaochang Li +2
Split Learning (SL) is an emerging privacy-preserving machine learning technique that enables resource constrained edge devices to participate in model training by partitioning a m…
RLPR: Extrapolating RLVR to General Domains without Verifiers
Tianyu Yu, Bo Ji, Shouli Wang +9
Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates promising potential in advancing the reasoning capabilities of LLMs. However, its success remains largely confine…