activity
20222024
most citedMCTformer+: Multi-Class Token Transformer for Weakly Supervised Semantic Segmentation

2 citations · 8 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV20241 cited

Motion-Oriented Compositional Neural Radiance Fields for Monocular Dynamic Human Modeling

Jaehyeok Kim, Dongyoon Wee, Dan Xu

This paper introduces Motion-oriented Compositional Neural Radiance Fields (MoCo-NeRF), a framework designed to perform free-viewpoint rendering of monocular human videos via novel…

cs.CV2024

Auxiliary Tasks Enhanced Dual-affinity Learning for Weakly Supervised Semantic Segmentation

Lian Xu, Mohammed Bennamoun, Farid Boussaid +3

Most existing weakly supervised semantic segmentation (WSSS) methods rely on Class Activation Mapping (CAM) to extract coarse class-specific localization maps using image-level lab…

cs.CV20232 cited

MCTformer+: Multi-Class Token Transformer for Weakly Supervised Semantic Segmentation

Lian Xu, Mohammed Bennamoun, Farid Boussaid +3

This paper proposes a novel transformer-based framework that aims to enhance weakly supervised semantic segmentation (WSSS) by generating accurate class-specific object localizatio…

cs.CV20232 cited

DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment

Lewei Yao, Jianhua Han, Xiaodan Liang +4

This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike…

cs.CV20231 cited

Joint 2D-3D Multi-Task Learning on Cityscapes-3D: 3D Detection, Segmentation, and Depth Estimation

Hanrong Ye, Dan Xu

This report serves as a supplementary document for TaskPrompter, detailing its implementation on a new joint 2D-3D multi-task learning benchmark based on Cityscapes-3D. TaskPrompte…

cs.CV20232 cited

You Only Train Once: Multi-Identity Free-Viewpoint Neural Human Rendering from Monocular Videos

Jaehyeok Kim, Dongyoon Wee, Dan Xu

We introduce You Only Train Once (YOTO), a dynamic human generation framework, which performs free-viewpoint rendering of different human identities with distinct motions, via only…