13 papers
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
Xiao Cai, Pengpeng Zeng, Ji Zhang +3
Precise spatial fidelity in Image-to-3D multi-instance generation is critical for downstream real-world applications. Recent work attempts to address this by fine-tuning pre-traine…
Reversible Inversion for Training-Free Exemplar-guided Image Editing
Yuke Li, Lianli Gao, Ji Zhang +5
Exemplar-guided Image Editing (EIE) aims to modify a source image according to a visual reference. Existing approaches often require large-scale pre-training to learn relationships…
Policy Contrastive Decoding for Robotic Foundation Models
Shihan Wu, Xu Luo, Ji Zhang +4
Robotic foundation models, or generalist robot policies, hold immense potential to enable flexible, general-purpose and dexterous robotic systems. Despite their advancements, our e…
Structure-aware Prompt Adaptation from Seen to Unseen for Open-Vocabulary Compositional Zero-Shot Learning
Yihang Duan, Jiong Wang, Pengpeng Zeng +5
The goal of Open-Vocabulary Compositional Zero-Shot Learning (OV-CZSL) is to recognize attribute-object compositions in the open-vocabulary setting, where compositions of both seen…
Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols
Xu Luo, Ji Zhang, Lianli Gao +2
Few-shot transfer has been revolutionized by stronger pre-trained models and improved adaptation algorithms.However, there lacks a unified, rigorous evaluation protocol that is bot…
Beyond the Majority: Long-tail Imitation Learning for Robotic Manipulation
Junhong Zhu, Ji Zhang, Jingkuan Song +2
While generalist robot policies hold significant promise for learning diverse manipulation skills through imitation, their performance is often hindered by the long-tail distributi…