Publications (326)
ImprovDML: Improved Trade-off in Private Byzantine-Resilient Distributed Machine Learning
Bing Liu, Chengcheng Zhao, Li Chai +2
Jointly addressing Byzantine attacks and privacy leakage in distributed machine learning (DML) has become an important issue. A common strategy involves integrating Byzantine-resil…
Tutel: Adaptive Mixture-of-Experts at Scale
Changho Hwang, Wei Cui, Yifan Xiong +12
Sparsely-gated mixture-of-experts (MoE) has been widely adopted to scale deep learning models to trillion-plus parameters with fixed computational cost. The algorithmic performance…
Robust Watermarks Leak: Channel-Aware Feature Extraction Enables Adversarial Watermark Manipulation
Zhongjie Ba, Yitao Zhang, Peng Cheng +4
Watermarking plays a key role in the provenance and detection of AI-generated content. While existing methods prioritize robustness against real-world distortions (e.g., JPEG compr…
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models
Zhenghao Lin, Zihao Tang, Xiao Liu +31
We introduce Sigma, an efficient large language model specialized for the system domain, empowered by a novel architecture including DiffQKV attention, and pre-trained on our metic…
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
Miao Li, Ming-Bin Chen, Bo Tang +8
We present NewsBench, a novel evaluation framework to systematically assess the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. Our c…
Synthesizing and characterization of hole doped nickel based superconductor (LaSr)NiAsO
Lei Fang, Huan Yang, Peng Cheng +3
We report the synthesizing and characterization of the hole doped Ni-based superconductor (. By substituting La with Sr, the superconducting transition temper…