most citedAudio-Visual Segmentation with Semantics

12 citations · 24 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2024

NeRFDeformer: NeRF Transformation from a Single View via 3D Scene Flows

Zhenggang Tang, Zhongzheng Ren, Xiaoming Zhao +4

We present a method for automatically modifying a NeRF representation based on a single observation of a non-rigid transformed version of the original scene. Our method defines the…

cs.CV20235 cited

Diff-DOPE: Differentiable Deep Object Pose Estimation

Jonathan Tremblay, Bowen Wen, Valts Blukis +3

We introduce Diff-DOPE, a 6-DoF pose refiner that takes as input an image, a 3D textured model of an object, and an initial pose of the object. The method uses differentiable rende…

cs.CV2023

TTA-COPE: Test-Time Adaptation for Category-Level Object Pose Estimation

Taeyeop Lee, Jonathan Tremblay, Valts Blukis +6

Test-time adaptation methods have been gaining attention recently as a practical solution for addressing source-to-target domain gaps by gradually updating the model without requir…

cs.CV20235 cited

BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects

Bowen Wen, Jonathan Tremblay, Valts Blukis +6

We present a near real-time method for 6-DoF tracking of an unknown object from a monocular RGBD video sequence, while simultaneously performing neural 3D reconstruction of the obj…

cs.CV20232 cited

Affordance Diffusion: Synthesizing Hand-Object Interactions

Yufei Ye, Xueting Li, Abhinav Gupta +5

Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for syn…

cs.CV202312 cited

Audio-Visual Segmentation with Semantics

Jinxing Zhou, Xuyang Shen, Jianyuan Wang +8

We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame…