Publications (7)
Benchmark Test-Time Scaling of General LLM Agents
Xiaochuan Li, Ryan Ming, Pranav Setlur +6
LLM agents are increasingly expected to function as general-purpose systems capable of resolving open-ended user requests. While existing benchmarks focus on domain-aware environme…
Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models
Annie S. Chen, Alec M. Lessing, Andy Tang +4
Legged robots are physically capable of navigating a diverse variety of environments and overcoming a wide range of obstructions. For example, in a search and rescue mission, a leg…
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
William Chen, Jagdeep Singh Bhatia, Catherine Glossop +6
Pretrained vision-language models (VLMs) can make semantic and visual inferences across diverse settings, providing valuable common-sense priors for robotic control. However, effec…
Improving Robotic Generalist Policies via Flow Reversal Steering
Andy Tang, William Chen, Andrew Wagenmaker +2
Generalist policies can learn a wide range of skills from diverse robot datasets. In order to solve or improve on challenging new tasks, we need a way to infer and invoke the appro…
Learning Long-Context Diffusion Policies via Past-Token Prediction
Marcel Torne, Andy Tang, Yuejiang Liu +1
Reasoning over long sequences of observations and actions is essential for many robotic tasks. Yet, learning effective long-context policies from demonstrations remains challenging…
Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
Prahaladh Chandrahasan, Jiahe Jin, Zhihan Zhang +9
Effectively evaluating deep research agents that autonomously search the web, analyze information, and generate reports remains a major challenge, particularly when it comes to ass…