2 papers
cs.AI2026
MARS: Margin-Adversarial Risk-controlled Stopping for Parallel LLM Test-time Scaling
Wenbo Chen, Puheng Li, Mengyang Liu +2
Parallel test-time scaling samples many reasoning traces and majority-votes their answers, improving LLM accuracy but requiring traces to run to completion, incurring substantial c…
cs.LG2025
AutoTailor: Automatic and Efficient Adaptive Model Deployment for Diverse Edge Devices
Mengyang Liu, Chenyu Lu, Haodong Tian +7
On-device machine learning (ML) has become a fundamental component of emerging mobile applications. Adaptive model deployment delivers efficient inference for heterogeneous device…