2 citations · 2 across the 1 of their papers we have counts for
4 papers
UniTok: A Unified Tokenizer for Visual Generation and Understanding
Chuofan Ma, Yi Jiang, Junfeng Wu +5
Visual generative and understanding models typically rely on distinct tokenizers to process images, presenting a key challenge for unifying them within a single framework. Recent s…
Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning
Runyu Ding, Yuzhe Qin, Jiyue Zhu +5
Teleoperation is a crucial tool for collecting human demonstrations, but controlling robots with bimanual dexterous hands remains a challenge. Existing teleoperation systems strugg…
Can 3D Vision-Language Models Truly Understand Natural Language?
Weipeng Deng, Jihan Yang, Runyu Ding +4
Rapid advancements in 3D vision-language (3D-VL) tasks have opened up new avenues for human interaction with embodied agents or robots using natural language. Despite this progress…
V-IRL: Grounding Virtual Intelligence in Real Life
Jihan Yang, Runyu Ding, Ellis Brown +2
There is a sensory gulf between the Earth that humans inhabit and the digital realms in which modern AI agents are created. To develop AI agents that can sense, think, and act as f…