activity
20212024
most citedAligning Large Multimodal Models with Factually Augmented RLHF

12 citations · 39 across the 14 of their papers we have counts for

collaborators

12 papers

cs.CV2024

TAMM: TriAdapter Multi-Modal Learning for 3D Shape Understanding

Zhihao Zhang, Shengcao Cao, Yu-Xiong Wang

The limited scale of current 3D shape datasets hinders the advancements in 3D shape understanding, and motivates multi-modal learning approaches which transfer learned knowledge fr…

cs.CV20244 cited

HASSOD: Hierarchical Adaptive Self-Supervised Object Detection

Shengcao Cao, Dhiraj Joshi, Liang-Yan Gui +1

The human visual perception system demonstrates exceptional capabilities in learning without explicit supervision and understanding the part-to-whole composition of objects. Drawin…

cs.CV20247 cited

ViCA-NeRF: View-Consistency-Aware 3D Editing of Neural Radiance Fields

Jiahua Dong, Yu-Xiong Wang

We introduce ViCA-NeRF, the first view-consistency-aware method for 3D editing with text instructions. In addition to the implicit neural radiance field (NeRF) modeling, our key in…

cs.CV2023

Improving Equivariance in State-of-the-Art Supervised Depth and Normal Predictors

Yuanyi Zhong, Anand Bhattad, Yu-Xiong Wang +1

Dense depth and surface normal predictors should possess the equivariant property to cropping-and-resizing -- cropping the input image should result in cropping the same output ima…

cs.CV2023

Streaming Motion Forecasting for Autonomous Driving

Ziqi Pang, Deva Ramanan, Mengtian Li +1

Trajectory forecasting is a widely-studied problem for autonomous navigation. However, existing benchmarks evaluate forecasting based on independent snapshots of trajectories, whic…

cs.CV2023

Multi-task View Synthesis with Neural Radiance Fields

Shuhong Zheng, Zhipeng Bao, Martial Hebert +1

Multi-task visual learning is a critical aspect of computer vision. Current research, however, predominantly concentrates on the multi-task dense prediction setting, which overlook…