62 citations · 69 across the 5 of their papers we have counts for
8 papers · 1 filter
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
Mu Chen, Liulei Li, Wenguan Wang +1
Top-leading solutions for Video Scene Graph Generation (VSGG) typically adopt an offline pipeline. Though demonstrating promising performance, they remain unable to handle real-tim…
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
Liulei Li, Wenguan Wang, Yi Yang
Prevalent human-object interaction (HOI) detection approaches typically leverage large-scale visual-linguistic models to help recognize events involving humans and objects. Though…
Vision-Language Navigation with Energy-Based Policy
Rui Liu, Wenguan Wang, Yi Yang
Vision-language navigation (VLN) requires an agent to execute actions following human instructions. Existing VLN models are optimized through expert demonstrations by supervised be…
Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph Generation
Minghan Chen, Guikun Chen, Wenguan Wang +1
DETR introduces a simplified one-stage framework for scene graph generation (SGG) but faces challenges of sparse supervision and false negative samples. The former occurs because e…
GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models
Chen Liang, Wenguan Wang, Jiaxu Miao +1
Prevalent semantic segmentation solutions are, in essence, a dense discriminative classifier of p(class|pixel feature). Though straightforward, this de facto paradigm neglects the…
Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence Learning
Liulei Li, Tianfei Zhou, Wenguan Wang +3
Our target is to learn visual correspondence from unlabeled videos. We develop LIIR, a locality-aware inter-and intra-video reconstruction framework that fills in three missing pie…