270 citations · 546 across the 44 of their papers we have counts for
34 papers · 1 filter
Semantic-SAM: Segment and Recognize Anything at Any Granularity
Feng Li, Hao Zhang, Peize Sun +6
In this paper, we introduce Semantic-SAM, a universal image segmentation model to enable segment and recognize anything at any desired granularity. Our model offers two key advanta…
LipsFormer: Introducing Lipschitz Continuity to Vision Transformers
Xianbiao Qi, Jianan Wang, Yihao Chen +2
We present a Lipschitz continuous Transformer, called LipsFormer, to pursue training stability both theoretically and empirically for Transformer-based models. In contrast to previ…
DisCo-CLIP: A Distributed Contrastive Loss for Memory Efficient CLIP Training
Yihao Chen, Xianbiao Qi, Jianan Wang +1
We propose DisCo-CLIP, a distributed memory-efficient CLIP training approach, to reduce the memory consumption of contrastive loss when training contrastive learning models. Our ap…
HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation
Xuan Ju, Ailing Zeng, Chenchen Zhao +3
Controllable human image generation (HIG) has numerous real-life applications. State-of-the-art solutions, such as ControlNet and T2I-Adapter, introduce an additional learnable bra…
Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes
Xuan Ju, Ailing Zeng, Jianan Wang +2
Humans have long been recorded in a variety of forms since antiquity. For example, sculptures and paintings were the primary media for depicting human beings before the invention o…
One-to-Few Label Assignment for End-to-End Dense Detection
Shuai Li, Minghan Li, Ruihuang Li +2
One-to-one (o2o) label assignment plays a key role for transformer based end-to-end detection, and it has been recently introduced in fully convolutional detectors for end-to-end d…