4 papers
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference
Junyi Luo, Xinting Jiang, Tai-Hao Wen +9
Microscaling (MX) is now the standard for low-bit large language model (LLM) inference. Its 4-bit form MXFP4 still loses substantial accuracy, because existing MX formats fix eithe…
LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models
Xinting Jiang, Junyi Luo, Ruichen Qi +4
Sub-billion-parameter Transformer language models are increasingly deployed on edge devices, where the privacy, latency, and operating-cost advantages of on-device inference are co…
Mitigating Classical Resource Costs in Quantum Error Correction via Generalized qLDPC Predecoding
Alexander Knapen, Junyi Luo, Guanchen Tao +6
Large-scale fault-tolerant quantum computing (FTQC) will require quantum-classical interfaces (QCIs) that orchestrate real-time decoding over thousands to millions of logical qubit…
Asymmetric stress engineering of dense dislocations in brittle superconductors for strong vortex pinning
Meng Han, Chiheng Dong, Chao Yao +17
Large lossless currents in high-temperature superconductors (HTS) critically rely on dense defects with suitable size and dimensionality to pin vortices, with dislocations being pa…