32 citations · 40 across the 3 of their papers we have counts for
7 papers
Visualizing and Understanding Patch Interactions in Vision Transformer
Jie Ma, Yalong Bai, Bineng Zhong +3
Vision Transformer (ViT) has become a leading tool in various computer vision tasks, owing to its unique self-attention mechanism that learns visual representations explicitly thro…
Freeform Body Motion Generation from Speech
Jing Xu, Wei Zhang, Yalong Bai +2
People naturally conduct spontaneous body motions to enhance their speeches while giving talks. Body motion generation from speech is inherently difficult due to the non-determinis…
Exploiting Relationship for Complex-scene Image Generation
Tianyu Hua, Hongdong Zheng, Yalong Bai +3
The significant progress on Generative Adversarial Networks (GANs) has facilitated realistic single-object image generation based on language input. However, complex-scene generati…
Products-10K: A Large-scale Product Recognition Dataset
Yalong Bai, Yuxiang Chen, Wei Yu +2
With the rapid development of electronic commerce, the way of shopping has experienced a revolutionary evolution. To fully meet customers' massive and diverse online shopping needs…
Look-into-Object: Self-supervised Structure Modeling for Object Recognition
Mohan Zhou, Yalong Bai, Wei Zhang +2
Most object recognition approaches predominantly focus on learning discriminative visual patterns while overlooking the holistic object structure. Though important, structure model…
Relationship-Aware Spatial Perception Fusion for Realistic Scene Layout Generation
Hongdong Zheng, Yalong Bai, Wei Zhang +1
The significant progress on Generative Adversarial Networks (GANs) have made it possible to generate surprisingly realistic images for single object based on natural language descr…