6 papers
Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference
Yuhang Gan, Yiwei Yang, Yuyi Li +6
Long-running LLM agents keep valuable state resident on GPUs: KV caches, request schedulers, communication state, and sometimes online adapters. Losing this state after a GPU or co…
One-Prompt Censorship Evasion via Generative Diffusion Models
Shiyi Ling, Yuhang Gan, Chen Qian
The escalating arms race between Internet censorship and evasion has driven censors to evolve from static rule-based filtering to sophisticated deep learning-based traffic analysis…
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
Fei Fang, Yifan Hua, Shengze Wang +4
While significant progress has been made in research and development on open-source and cost-efficient large-language models (LLMs), serving scalability remains a critical challeng…
Scalable Community Detection Using Quantum Hamiltonian Descent and QUBO Formulation
Jinglei Cheng, Ruilin Zhou, Yuhang Gan +2
We present a quantum-inspired algorithm that utilizes Quantum Hamiltonian Descent (QHD) for efficient community detection. Our approach reformulates the community detection task as…
Optimizing Compilation for Distributed Quantum Computing via Clustering and Annealing
Ruilin Zhou, Jinglei Cheng, Yuhang Gan +2
Efficiently mapping quantum programs onto Distributed quantum computing (DQC) are challenging, particularly when considering the heterogeneous quantum processing units (QPUs) with…
CloudQC: A Network-aware Framework for Multi-tenant Distributed Quantum Computing
Ruilin Zhou, Yuhang Gan, Yi Liu +1
Distributed quantum computing (DQC) that allows a large quantum circuit to be executed simultaneously on multiple quantum processing units (QPUs) becomes a promising approach to in…