29 citations · 96 across the 38 of their papers we have counts for
22 papers
3D Face Modeling via Weakly-supervised Disentanglement Network joint Identity-consistency Prior
Guohao Li, Hongyu Yang, Di Huang +1
Generative 3D face models featuring disentangled controlling factors hold immense potential for diverse applications in computer vision and computer graphics. However, previous 3D…
iVPT: Improving Task-relevant Information Sharing in Visual Prompt Tuning by Cross-layer Dynamic Connection
Nan Zhou, Jiaxin Chen, Di Huang
Recent progress has shown great potential of visual prompt tuning (VPT) when adapting pre-trained vision transformers to various downstream tasks. However, most existing solutions…
Generalizing 6-DoF Grasp Detection via Domain Prior Knowledge
Haoxiang Ma, Modi Shi, Boyang Gao +1
We focus on the generalization ability of the 6-DoF grasp detection method in this paper. While learning-based grasp detection methods can predict grasp poses for unseen objects us…
Emergent Communication for Rules Reasoning
Yuxuan Guo, Yifan Hao, Rui Zhang +14
Research on emergent communication between deep-learning-based agents has received extensive attention due to its inspiration for linguistics and artificial intelligence. However,…
BirdSAT: Cross-View Contrastive Masked Autoencoders for Bird Species Classification and Mapping
Srikumar Sastry, Subash Khanal, Aayush Dhakal +2
We propose a metadata-aware self-supervised learning~(SSL)~framework useful for fine-grained classification and ecological mapping of bird species around the world. Our framework u…
Self-driven Grounding: Large Language Model Agents with Automatical Language-aligned Skill Learning
Shaohui Peng, Xing Hu, Qi Yi +9
Large language models (LLMs) show their powerful automatic reasoning and planning capability with a wealth of semantic knowledge about the human world. However, the grounding probl…