Publications (28)
Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization
Ziyuan Tang, Tianshi Xu, Yousef Saad +1
Muon-type optimizers construct update directions for dense neural-network weights by applying a finite Newton-Schulz map to momentum-gradient matrices. For an matrix,…
FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference
Chenqi Lin, Tianshi Xu, Zebin Yang +3
With the fast evolution of large language models (LLMs), privacy concerns with user queries arise as they may contain sensitive information. Private inference based on homomorphic…
Falcon: Accelerating Homomorphically Encrypted Convolutions for Efficient Private Mobile Network Inference
Tianshi Xu, Meng Li, Runsheng Wang +1
Efficient networks, e.g., MobileNetV2, EfficientNet, etc, achieves state-of-the-art (SOTA) accuracy with lightweight computation. However, existing homomorphic encryption (HE)-base…
Factored Sparse Approximate Inverse Preconditioning via Spectral Optimization
Francesco Brarda, Tianshi Xu, Vassilis Kalantzis +2
In this paper, we study value selection for fixed-pattern factorized sparse approximate inverse preconditioners. Given a prescribed sparsity pattern for a factor we choose its…
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
Tianshi Xu, Huifeng Wen, Meng Li
LLM agents are shaped not only by their language models, but also by the runtime harness that mediates observation, tool use, action execution, feedback interpretation, and traject…
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
Chenqi Lin, Kang Yang, Tianshi Xu +6
With the wide application of machine learning (ML), privacy concerns arise with user data as they may contain sensitive information. Privacy-preserving ML (PPML) based on cryptogra…