5 papers
Hierarchical Policy Learning via Spectral Decomposition
Shuxin Cao, Liquan Wang, Walker Byrnes +3
In this paper, we identify a semantic decomposition in robot action sequences, separating task-level motion intent from execution-level refinements. By analyzing actions in the spe…
VISTA: Enhancing Visual Conditioning via Track-Following Preference Optimization in Vision-Language-Action Models
Yiye Chen, Yanan Jian, Xiaoyi Dong +5
Vision-Language-Action (VLA) models have demonstrated strong performance across a wide range of robotic manipulation tasks. Despite the success, extending large pretrained Vision-L…
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
Jing Wu, Daphne Barretto, Yiye Chen +4
Long-horizon, repetitive workflows are common in professional settings, such as processing expense reports from receipts and entering student grades from exam papers. These tasks a…
Schema-Guided Scene-Graph Reasoning based on Multi-Agent Large Language Model System
Yiye Chen, Harpreet Sawhney, Nicholas Gydé +4
Scene graphs have emerged as a structured and serializable environment representation for grounded spatial reasoning with Large Language Models (LLMs). In this work, we propose SG^…
GASP: Gaussian Avatars with Synthetic Priors
Jack Saunders, Charlie Hewitt, Yanan Jian +8
Gaussian Splatting has changed the game for real-time photo-realistic rendering. One of the most popular applications of Gaussian Splatting is to create animatable avatars, known a…