29 citations · 114 across the 21 of their papers we have counts for
13 papers · 1 filter
Multi-modal Prompting for Low-Shot Temporal Action Localization
Chen Ju, Zeqian Li, Peisen Zhao +5
In this paper, we consider the problem of temporal action localization under low-shot (zero-shot & few-shot) scenario, with the goal of detecting and classifying the action instanc…
DR2: Diffusion-based Robust Degradation Remover for Blind Face Restoration
Zhixin Wang, Xiaoyun Zhang, Ziying Zhang +4
Blind face restoration usually synthesizes degraded low-quality data with a pre-defined degradation model for training, while more complex cases could happen in the real world. Thi…
Boundary-aware Supervoxel-level Iteratively Refined Interactive 3D Image Segmentation with Multi-agent Reinforcement Learning
Chaofan Ma, Qisen Xu, Xiangfeng Wang +4
Interactive segmentation has recently been explored to effectively and efficiently harvest high-quality segmentation masks by iteratively incorporating user hints. While iterative…
DiffusionSeg: Adapting Diffusion Towards Unsupervised Object Discovery
Chaofan Ma, Yuhuan Yang, Chen Ju +5
Learning from a large corpus of data, pre-trained models have achieved impressive progress nowadays. As popular generative pre-training, diffusion models capture both low-level vis…
Controllable Mesh Generation Through Sparse Latent Point Diffusion Models
Zhaoyang Lyu, Jinyi Wang, Yuwei An +3
Mesh generation is of great value in various applications involving computer graphics and virtual content, yet designing generative models for meshes is challenging due to their ir…
PMC-CLIP: Contrastive Language-Image Pre-training using Biomedical Documents
Weixiong Lin, Ziheng Zhao, Xiaoman Zhang +4
Foundation models trained on large-scale dataset gain a recent surge in CV and NLP. In contrast, development in biomedical domain lags far behind due to data scarcity. To address t…