16 papers
Thought-Level Beam Search for Reasoning
Lijie Yang, Hongyin Luo, Jiawei Zhao +2
Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question fr…
Wherefore Art Thou? Provenance-Guided Automatic Online Debugging with Lumos
Jingyuan Chen, Lei Zhang, Leon Schuermann +3
Debugging distributed systems in-production is inevitable and hard. Myriad interactions between concurrent components in modern, complex and large-scale systems cause non-determini…
Skim: Speculative Execution for Fast and Efficient Web Agents
Mike Wong, Kevin Hsieh, Suman Nath +1
Skim is a speculative execution framework for web agents that exploits the predictable structure of purpose-built websites. Today's web-agent expense is not intrinsic to the tasks…
Kairos: A Scalable Serving System for Physical AI
Yinwei Dai, Ganesh Ananthanarayanan, Landon Cox +3
Physical AI is experiencing rapid growth with frontier foundation models increasing its capabilities across general environments. Physical AI tasks are characterized by inference p…
Geometry Guided Self-Consistency for Physical AI
Yinwei Dai, Zhuofu Chen, Lijie Yang +1
State-of-the-art physical AI models generate a chunk of actions per inference through diffusion or flow matching, iteratively refining an initial noise sample into an action trajec…
Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents
Zhuofu Chen, Rui Pan, Yinwei Dai +1
To cope with the large contexts that long-horizon LLM agents produce, modern frameworks increasingly rely on compaction -- invoking an LLM to rewrite the accumulated trajectory int…