most citedNeural Textured Deformable Meshes for Robust Analysis-by-Synthesis

1 citations · 2 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2024

Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data

Wufei Ma, Kai Li, Zhongshi Jiang +5

Recent video-text foundation models have demonstrated strong performance on a wide variety of downstream video understanding tasks. Can these video-text models genuinely understand…

cs.CV20241 cited

ImageNet3D: Towards General-Purpose Object-Level 3D Understanding

Wufei Ma, Guanning Zeng, Guofeng Zhang +5

A vision model with general-purpose object-level 3D understanding should be capable of inferring both 2D (e.g., class name and bounding box) and 3D information (e.g., 3D location a…

cs.CV2024

Uncertainty-Aware Deep Video Compression with Ensembles

Wufei Ma, Jiahao Li, Bin Li +1

Deep learning-based video compression is a challenging task, and many previous state-of-the-art learning-based video codecs use optical flows to exploit the temporal correlation be…

cs.CV2023

3D-Aware Visual Question Answering about Parts, Poses and Occlusions

Xingrui Wang, Wufei Ma, Zhuowan Li +2

Despite rapid progress in Visual question answering (VQA), existing datasets and models mainly focus on testing reasoning in 2D. However, it is important that VQA models also under…

cs.CV20231 cited

Neural Textured Deformable Meshes for Robust Analysis-by-Synthesis

Angtian Wang, Wufei Ma, Alan Yuille +1

Human vision demonstrates higher robustness than current AI algorithms under out-of-distribution scenarios. It has been conjectured such robustness benefits from performing analysi…

cs.CV2023

Robust Category-Level 3D Pose Estimation from Synthetic Data

Jiahao Yang, Wufei Ma, Angtian Wang +3

Obtaining accurate 3D object poses is vital for numerous computer vision applications, such as 3D reconstruction and scene understanding. However, annotating real-world objects is…