papers

Publications (28)

math.NA2026

Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization

Ziyuan Tang, Tianshi Xu, Yousef Saad +1

Muon-type optimizers construct update directions for dense neural-network weights by applying a finite Newton-Schulz map to momentum-gradient matrices. For an matrix,…

cs.CR2024

FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference

Chenqi Lin, Tianshi Xu, Zebin Yang +3

With the fast evolution of large language models (LLMs), privacy concerns with user queries arise as they may contain sensitive information. Private inference based on homomorphic…

cs.CR2023

Falcon: Accelerating Homomorphically Encrypted Convolutions for Efficient Private Mobile Network Inference

Tianshi Xu, Meng Li, Runsheng Wang +1

Efficient networks, e.g., MobileNetV2, EfficientNet, etc, achieves state-of-the-art (SOTA) accuracy with lightweight computation. However, existing homomorphic encryption (HE)-base…

math.NA2026

Factored Sparse Approximate Inverse Preconditioning via Spectral Optimization

Francesco Brarda, Tianshi Xu, Vassilis Kalantzis +2

In this paper, we study value selection for fixed-pattern factorized sparse approximate inverse preconditioners. Given a prescribed sparsity pattern for a factor we choose its…

cs.AI2026

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents

Tianshi Xu, Huifeng Wen, Meng Li

LLM agents are shaped not only by their language models, but also by the runtime harness that mediates observation, tool use, action execution, feedback interpretation, and traject…

cs.AR2025

Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing

Chenqi Lin, Kang Yang, Tianshi Xu +6

With the wide application of machine learning (ML), privacy concerns arise with user data as they may contain sensitive information. Privacy-preserving ML (PPML) based on cryptogra…