6 papers
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
Ruiping Liu, Junwei Zheng, Yufan Chen +7
Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and obs…
InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing
Yebin Yang, Di Wen, Lei Qi +10
Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity…
FlowNar: Scalable Streaming Narration for Long-Form Videos
Zeyun Zhong, Manuel Martin, Chengzhi Wu +4
Recent Large Multimodal Models (LMMs), primarily designed for offline settings, are ill-suited for the dynamic requirements of streaming video. While recent online adaptations impr…
6D Pose Estimation on Point Cloud Data through Prior Knowledge Integration: A Case Study in Autonomous Disassembly
Chengzhi Wu, Hao Fu, Jan-Philipp Kaiser +5
The accurate estimation of 6D pose remains a challenging task within the computer vision domain, even when utilizing 3D point cloud data. Conversely, in the manufacturing domain, i…
A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning
Chengzhi Wu, Qianliang Huang, Kun Jin +2
Contrastive learning is an essential method in self-supervised learning. It primarily employs a multi-branch strategy to compare latent representations obtained from different bran…
SAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global Uniformity
Chengzhi Wu, Yuxin Wan, Hao Fu +5
Driven by the increasing demand for accurate and efficient representation of 3D data in various domains, point cloud sampling has emerged as a pivotal research topic in 3D computer…