papers

Publications (11)

cs.LG2025

MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators

Beichen Huang, Yueming Yuan, Zelei Shao +1

A critical approach for efficiently deploying Mixture-of-Experts (MoE) models with massive parameters is quantization. However, state-of-the-art MoE models suffer from non-negligib…

cs.AI2026

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

Zhixiang Liang, Yifei Liu, Yidan Huang +5

Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long…

cs.SE2026

Persistent Recursive Worlds Enable Autonomous Software Evolution

Beichen Huang, Zhenyu Liang, Bowen Zheng +1

Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessi…

cs.LG2026

Hidden States as Early Signals: Step-level Trace Evaluation and Pruning for Efficient Test-Time Scaling

Zhixiang Liang, Beichen Huang, Zheng Wang +1

Large Language Models (LLMs) can enhance reasoning capabilities through test-time scaling by generating multiple traces. However, the combination of lengthy reasoning traces with m…

cs.NE2024

Exploring the True Potential: Evaluating the Black-box Optimization Capability of Large Language Models

Beichen Huang, Xingyu Wu, Yu Zhou +4

Large language models (LLMs) have demonstrated exceptional performance not only in natural language processing tasks but also in a great variety of non-linguistic domains. In diver…

physics.atom-ph2024

Proposal for the generation of continuous-wave vacuum ultraviolet laser light for Th-229 isomer precision spectroscopy

Qi Xiao, Gleb Penyazkov, Ruihan Yu +6

We propose to generate continuous-wave vacuum ultraviolet (VUV) laser light at 148.4 nm using four-wave mixing in cadmium vapor for precision spectroscopy of the Th-229 isomer tran…