24 citations · 51 across the 10 of their papers we have counts for
10 papers
Visual Commonsense based Heterogeneous Graph Contrastive Learning
Zongzhao Li, Xiangyu Zhu, Xi Zhang +2
How to select relevant key objects and reason about the complex relationships cross vision and linguistic domain are two key issues in many multi-modality applications such as visu…
DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action Localization
Xiaojun Tang, Junsong Fan, Chuanchen Luo +3
Weakly-supervised temporal action localization (WTAL) is a practical yet challenging task. Due to large-scale datasets, most existing methods use a network pretrained in other data…
Visually-Guided Sound Source Separation with Audio-Visual Predictive Coding
Zengjie Song, Zhaoxiang Zhang
The framework of visually-guided sound source separation generally consists of three parts: visual feature extraction, multimodal feature fusion, and sound signal processing. An on…
Tracking Objects with 3D Representation from Videos
Jiawei He, Lue Fan, Yuqi Wang +4
Data association is a knotty problem for 2D Multiple Object Tracking due to the object occlusion. However, in 3D space, data association is not so hard. Only with a 3D Kalman Filte…
BMAD: Benchmarks for Medical Anomaly Detection
Jinan Bao, Hanshi Sun, Hanqiu Deng +3
Anomaly detection (AD) is a fundamental research problem in machine learning and computer vision, with practical applications in industrial inspection, video surveillance, and medi…
Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory
Xizhou Zhu, Yuntao Chen, Hao Tian +10
The captivating realm of Minecraft has attracted substantial research interest in recent years, serving as a rich platform for developing intelligent agents capable of functioning…