7 papers
Self-Supervised Skill Optimization
Siran Peng, Cuiyu Yang, Tianyu Fu +9
Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized with ground-truth (GT) feed…
Dive Into the Implicit Biases of Low-rank Vision-language Alignment
Mingjia Shi, Shuo Wang, Xiaobo Wang +7
Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates.…
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
Wei Dong, Tianyu Fu, Zhe Yu +9
As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion…
Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation
Jiahao Xu, Peiyuan Wang, Hanzhuo Zhang +7
In robotic manipulation, the tight coupling between grasping and motion planning often obscures the true source of failure, leading to inefficient trial-and-error. To enable effici…
UPA: Unsupervised Prompt Agent via Tree-Based Search and Selection
Siran Peng, Weisong Zhao, Tianyu Fu +6
Prompt agents have recently emerged as a promising paradigm for automated prompt optimization, framing prompt discovery as a sequential decision-making problem over a structured pr…
Mano Technical Report
Tianyu Fu, Anyang Su, Chenxu Zhao +20
Graphical user interfaces (GUIs) are the primary medium for human-computer interaction, yet automating GUI interactions remains challenging due to the complexity of visual elements…