29 citations · 40 across the 8 of their papers we have counts for
6 papers · 1 filter
What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception
Wanfang Su, Lixing Chen, Yang Bai +4
Multi-agent perception (MAP) allows autonomous systems to understand complex environments by interpreting data from multiple sources. This paper investigates intermediate collabora…
ConRF: Zero-shot Stylization of 3D Scenes with Conditioned Radiation Fields
Xingyu Miao, Yang Bai, Haoran Duan +4
Most of the existing works on arbitrary 3D NeRF style transfer required retraining on each single style condition. This work aims to achieve zero-shot controlled stylization in 3D…
DS-Depth: Dynamic and Static Depth Estimation via a Fusion Cost Volume
Xingyu Miao, Yang Bai, Haoran Duan +5
Self-supervised monocular depth estimation methods typically rely on the reprojection error to capture geometric relationships between successive frames in static environments. How…
Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations
Yiwu Zhong, Licheng Yu, Yang Bai +3
The abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn vi…
Temporal Segment Transformer for Action Segmentation
Zhichao Liu, Leshan Wang, Desen Zhou +5
Recognizing human actions from untrimmed videos is an important task in activity understanding, and poses unique challenges in modeling long-range temporal relations. Recent works…
Action Quality Assessment with Temporal Parsing Transformer
Yang Bai, Desen Zhou, Songyang Zhang +5
Action Quality Assessment(AQA) is important for action understanding and resolving the task poses unique challenges due to subtle visual differences. Existing state-of-the-art meth…