26 citations · 45 across the 3 of their papers we have counts for
4 papers
CgT-GAN: CLIP-guided Text GAN for Image Captioning
Jiarui Yu, Haoran Li, Yanbin Hao +3
The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated…
Bi-Touch: Bimanual Tactile Manipulation with Sim-to-Real Deep Reinforcement Learning
Yijiong Lin, Alex Church, Max Yang +4
Bimanual manipulation with tactile feedback will be key to human-level robot dexterity. However, this topic is less explored than single-arm settings, partly due to the availabilit…
Differentially Private Federated Knowledge Graphs Embedding
Hao Peng, Haoran Li, Yangqiu Song +2
Knowledge graph embedding plays an important role in knowledge representation, reasoning, and data mining applications. However, for multiple cross-domain knowledge graphs, state-o…
Music-oriented Dance Video Synthesis with Pose Perceptual Loss
Xuanchi Ren, Haoran Li, Zijian Huang +1
We present a learning-based approach with pose perceptual loss for automatic music video generation. Our method can produce a realistic dance video that conforms to the beats and r…