6 papers
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Yusen Feng, Bingchen Han, Jiangran Lyu +13
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fin…
ComUICoder: Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Element-wise Feedback
Jingyu Xiao, Jiantong Qin, Shuoqi Li +5
Multimodal Large Language Models (MLLMs) have demonstrated strong performance on the UI-to-code task, which aims to generate UI code from design mock-ups. However, when applied to…
FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real
Weiheng Liu, Yuxuan Wan, Jilong Wang +7
Generalizable object fetching in cluttered scenes remains a fundamental and application-critical challenge in embodied AI. Closely packed objects cause inevitable occlusions, makin…
Learning a Pessimistic Reward Model in RLHF
Yinglun Xu, Hangoo Kang, Tarun Suresh +2
This work proposes `PET', a novel pessimistic reward fine-tuning method, to learn a pessimistic reward model robust against reward hacking in offline reinforcement learning from hu…
SCANet: Correcting LEGO Assembly Errors with Self-Correct Assembly Network
Yuxuan Wan, Kaichen Zhou, jinhong Chen +1
Autonomous assembly in robotics and 3D vision presents significant challenges, particularly in ensuring assembly correctness. Presently, predominant methods such as MEPNet focus on…
DPStyler: Dynamic PromptStyler for Source-Free Domain Generalization
Yunlong Tang, Yuxuan Wan, Lei Qi +1
Source-Free Domain Generalization (SFDG) aims to develop a model that works for unseen target domains without relying on any source domain. Research in SFDG primarily bulids upon t…