activity
20202024
most citedDetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection

64 citations · 139 across the 27 of their papers we have counts for

collaborators
Showing cs.CVShow all

31 papers · 1 filter

cs.CV2024

HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models

Runhui Huang, Xinpeng Ding, Chunwei Wang +7

High-resolution inputs enable Large Vision-Language Models (LVLMs) to discern finer visual details, enhancing their comprehension capabilities. To reduce the training and computati…

cs.CV20241 cited

DetCLIPv3: Towards Versatile Generative Open-vocabulary Object Detection

Lewei Yao, Renjie Pi, Jianhua Han +5

Existing open-vocabulary object detectors typically require a predefined set of categories from users, significantly confining their application scenarios. In this paper, we introd…

cs.CV2024

LayerDiff: Exploring Text-guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model

Runhui Huang, Kaixin Cai, Jianhua Han +6

Despite the success of generating high-quality images given any text prompts by diffusion-based generative models, prior works directly generate the entire images, but cannot provi…

cs.CV2024

GS-CLIP: Gaussian Splatting for Contrastive Language-Image-3D Pretraining from Real-World Data

Haoyuan Li, Yanpeng Zhou, Yihan Zeng +2

3D Shape represented as point cloud has achieve advancements in multimodal pre-training to align image and language descriptions, which is curial to object identification, classifi…

cs.CV20245 cited

Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models

Xinpeng Ding, Jinahua Han, Hang Xu +3

The rise of multimodal large language models (MLLMs) has spurred interest in language-based driving tasks. However, existing research typically focuses on limited tasks and often o…

cs.CV2023

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

Cong Wang, Jiaxi Gu, Panwen Hu +3

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided i…