16 citations · 50 across the 6 of their papers we have counts for
7 papers
OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models
Jinze Bai, Rui Men, Hao Yang +15
Generalist models, which are capable of performing diverse multi-modal tasks in a task-agnostic way within a single model, have been explored recently. Being, hopefully, an alterna…
Pretrained Diffusion Models for Unified Human Motion Synthesis
Jianxin Ma, Shuai Bai, Chang Zhou
Generative modeling of human motion has broad applications in computer animation, virtual reality, and robotics. Conventional approaches develop separate models for different motio…
M6-Fashion: High-Fidelity Multi-modal Image Generation and Editing
Zhikang Li, Huiling Zhou, Shuai Bai +3
The fashion industry has diverse applications in multi-modal image generation and editing. It aims to create a desired high-fidelity image with the multi-modal conditional signal a…
Connecting Language and Vision for Natural Language-Based Vehicle Retrieval
Shuai Bai, Zhedong Zheng, Xiaohan Wang +5
Vehicle search is one basic task for the efficient traffic management in terms of the AI City. Most existing practices focus on the image-based vehicle matching, including vehicle…
Dense Relation Distillation with Context-aware Aggregation for Few-Shot Object Detection
Hanzhe Hu, Shuai Bai, Aoxue Li +2
Conventional deep learning based methods for object detection require a large amount of bounding box annotations for training, which is expensive to obtain such high quality annota…
Class-wise Dynamic Graph Convolution for Semantic Segmentation
Hanzhe Hu, Deyi Ji, Weihao Gan +3
Recent works have made great progress in semantic segmentation by exploiting contextual information in a local or global manner with dilated convolutions, pyramid pooling or self-a…