5 papers
DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA
Kango Yanagida, Tatsuya Aoki, Yuichiro Yoshikawa +1
Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category…
Relational Knowledge Distillation Brings DNN Representations Close Enough to Humans to Be Aligned Without Supervision
Yuria Shimizu, Soh Takahashi, Takato Horii +1
Linking the internal representations of deep neural networks (DNNs) to human mental representations is important for using DNNs as computational models of human vision. Existing DN…
VENOM: Versatile Embodied Network for Omni-bodied Motion tracking
Siddharth Padmanabhan, Kazuki Miyazawa, Takato Horii
Achieving expert-level expressive full-body motion tracking across multiple humanoids solely from demonstration data remains a challenging and relatively an underexplored problem i…
Public Evaluation on Potential Social Impacts of Fully Autonomous Cybernetic Avatars for Physical Support in Daily-Life Environments: Large-Scale Demonstration and Survey at Avatar Land
Lotfi El Hafi, Kazuma Onishi, Shoichi Hasegawa +18
Cybernetic avatars (CAs) are key components of an avatar-symbiotic society, enabling individuals to overcome physical limitations through virtual agents and robotic assistants. Whi…
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
Haruka Asanuma, Naoko Koide-Majima, Ken Nakamura +3
Recent studies have revealed that human emotions exhibit a high-dimensional, complex structure. A full capturing of this complexity requires new approaches, as conventional models…