320 citations · 572 across the 8 of their papers we have counts for
9 papers
DreamMesh: Jointly Manipulating and Texturing Triangle Meshes for Text-to-3D Generation
Haibo Yang, Yang Chen, Yingwei Pan +5
Learning radiance fields (NeRF) with powerful 2D diffusion models has garnered popularity for text-to-3D generation. Nevertheless, the implicit 3D representations of NeRF lack expl…
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks
Junke Wang, Dongdong Chen, Zuxuan Wu +7
This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based v…
Semi-Supervised Single-View 3D Reconstruction via Prototype Shape Priors
Zhen Xing, Hengduo Li, Zuxuan Wu +1
The performance of existing single-view 3D reconstruction methods heavily relies on large-scale 3D annotations. However, such annotations are tedious and expensive to collect. Semi…
ObjectFormer for Image Manipulation Detection and Localization
Junke Wang, Zuxuan Wu, Jingjing Chen +4
Recent advances in image editing techniques have posed serious challenges to the trustworthiness of multimedia data, which drives the research of image tampering detection. In this…
VideoLT: Large-scale Long-tailed Video Recognition
Xing Zhang, Zuxuan Wu, Zejia Weng +4
Label distributions in real-world are oftentimes long-tailed and imbalanced, resulting in biased models towards dominant labels. While long-tailed recognition has been extensively…
LiteEval: A Coarse-to-Fine Framework for Resource Efficient Video Recognition
Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang +1
This paper presents LiteEval, a simple yet effective coarse-to-fine framework for resource efficient video recognition, suitable for both online and offline scenarios. Exploiting d…