2 papers
cs.LG2026
Communication-Efficient Verifiable Attention for LLM Inference
Ziqun Chen, Ming Wu, Michael Heinrich +4
Computation integrity of remote large language model (LLM) serving can be questionable. For conventional deep neural networks (DNNs), the existing TEE-shielded DNN partitioning (TS…
cs.LG2024
Merit-based Fair Combinatorial Semi-Bandit with Unrestricted Feedback Delays
Ziqun Chen, Kechao Cai, Zhuoyue Chen +2
We study the stochastic combinatorial semi-bandit problem with unrestricted feedback delays under merit-based fairness constraints. This is motivated by applications such as crowds…