Publications (11)
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
Beichen Huang, Yueming Yuan, Zelei Shao +1
A critical approach for efficiently deploying Mixture-of-Experts (MoE) models with massive parameters is quantization. However, state-of-the-art MoE models suffer from non-negligib…
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
Zhixiang Liang, Yifei Liu, Yidan Huang +5
Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long…
Persistent Recursive Worlds Enable Autonomous Software Evolution
Beichen Huang, Zhenyu Liang, Bowen Zheng +1
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessi…
Hidden States as Early Signals: Step-level Trace Evaluation and Pruning for Efficient Test-Time Scaling
Zhixiang Liang, Beichen Huang, Zheng Wang +1
Large Language Models (LLMs) can enhance reasoning capabilities through test-time scaling by generating multiple traces. However, the combination of lengthy reasoning traces with m…
Exploring the True Potential: Evaluating the Black-box Optimization Capability of Large Language Models
Beichen Huang, Xingyu Wu, Yu Zhou +4
Large language models (LLMs) have demonstrated exceptional performance not only in natural language processing tasks but also in a great variety of non-linguistic domains. In diver…
Proposal for the generation of continuous-wave vacuum ultraviolet laser light for Th-229 isomer precision spectroscopy
Qi Xiao, Gleb Penyazkov, Ruihan Yu +6
We propose to generate continuous-wave vacuum ultraviolet (VUV) laser light at 148.4 nm using four-wave mixing in cadmium vapor for precision spectroscopy of the Th-229 isomer tran…