5 papers
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances
Jihoon Oh, Kento Kawaharazuka, Kei Okada
Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One prom…
MEVION: Low-Cost Open-Source Data Collection System for Powerful and High-Speed Dual-Arm Manipulation
Kento Kawaharazuka, Yoshiki Obinata, Hirokazu Ishida +6
The global competition for developing robotic foundation models is intensifying. Among the data collection systems used for dual-arm robots, ALOHA is representative of being low-co…
Towards Trustworthy LLM-Based Recommendation via Rationale Integration
Chung Park, Taesan Kim, Hyeongjun Yun +7
Traditional recommender systems (RS) have been primarily optimized for accuracy and short-term engagement, often overlooking transparency and trustworthiness. Recently, platforms s…
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
Kento Kawaharazuka, Jihoon Oh, Jun Yamada +2
Amid growing efforts to leverage advances in large language models (LLMs) and vision-language models (VLMs) for robotics, Vision-Language-Action (VLA) models have recently gained s…
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…