2 citations · 2 across the 7 of their papers we have counts for
8 papers
AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning
Zhiyue Zhao, Jingyi Wu, Hairuo Liu +5
Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spac…
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Haokai Zhang, Yuhang Ding, Yunshu Zhou +5
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing wor…
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
Liyang Li, Muzhi Zhu, Zhiyue Zhao +5
Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has largely been studied as passiv…
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
Liyang Li, Wen Wang, Canyu Zhao +4
Recent advances in Diffusion Transformers (DiTs) have enabled high-quality joint audio-video generation, producing videos with synchronized audio within a single model. However, ex…
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
Canyu Zhao, Xiaoman Li, Tianjian Feng +3
We introduce Tinker, a versatile framework for high-fidelity 3D editing that operates in both one-shot and few-shot regimes without any per-scene finetuning. Unlike prior technique…
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
Canyu Zhao, Yanlong Sun, Mingyu Liu +6
This paper's primary objective is to develop a robust generalist perception model capable of addressing multiple tasks under constraints of computational resources and limited trai…