3 papers
cs.CL2026
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
Zihan Qiu, Zekun Wang, Xiao Li +33
We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n…
cs.CR2026
On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference
Zhengyi Li, Yakai Wang, Kang Yang +6
For Transformer models, cryptographically secure inference ensures that the client learns only the final output, while the server learns nothing about the client's input. However,…
cs.DC2023
DistSim: A performance model of large-scale hybrid distributed DNN training
Guandong Lu, Runzhe Chen, Yakai Wang +8
With the ever-increasing computational demand of DNN training workloads, distributed training has been widely adopted. A combination of data, model and pipeline parallelism strateg…