activity
20162024
most citedCenterNet: Keypoint Triplets for Object Detection

168 citations · 1k across the 66 of their papers we have counts for

collaborators
Showing cs.CVShow all

93 papers · 1 filter

cs.CV2024

SAM-CP: Marrying SAM with Composable Prompts for Versatile Segmentation

Pengfei Chen, Lingxi Xie, Xinyue Huo +5

The Segment Anything model (SAM) has shown a generalized ability to group image pixels into patches, but applying it to semantic-aware segmentation still faces major challenges. Th…

cs.CV2024★ 2 cited

GaussianDreamerPro: Text to Manipulable 3D Gaussians with Highly Enhanced Quality

Taoran Yi, Jiemin Fang, Zanwei Zhou +7

Recently, 3D Gaussian splatting (3D-GS) has achieved great success in reconstructing and rendering real-world scenes. To transfer the high rendering quality to generation tasks, a…

cs.CV2024

Text-Animator: Controllable Visual Text Video Generation

Lin Liu, Quande Liu, Shengju Qian +5

Video generation is a challenging yet pivotal task in various industries, such as gaming, e-commerce, and advertising. One significant unresolved aspect within T2V is the effective…

cs.CV2024

Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation

Zelin Peng, Zhengqin Xu, Zhilin Zeng +2

Open-vocabulary semantic segmentation seeks to label each pixel in an image with arbitrary text descriptions. Vision-language foundation models, especially CLIP, have recently emer…

cs.CV2024★ 5 cited

AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation

Jiannan Ge, Lingxi Xie, Hongtao Xie +4

A serious issue that harms the performance of zero-shot visual recognition is named objective misalignment, i.e., the learning objective prioritizes improving the recognition accur…

cs.CV2024★ 3 cited

GaussianObject: High-Quality 3D Object Reconstruction from Four Views with Gaussian Splatting

Chen Yang, Sikuang Li, Jiemin Fang +5

Reconstructing and rendering 3D objects from highly sparse views is of critical importance for promoting applications of 3D vision techniques and improving user experience. However…