activity
20212026
most citedLocality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence Learning

2 citations · 2 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection

Xinhao Cai, Liulei Li, Gensheng Pei +3

Object detection in remote sensing images (RSIs) is challenged by the coexistence of geometric and spatial complexity: targets may appear with diverse aspect ratios, while spanning…

cs.CV2025

Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis

Xinhao Cai, Liulei Li, Gensheng Pei +4

This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive g…

cs.CV2025

DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation

Mu Chen, Liulei Li, Wenguan Wang +1

Top-leading solutions for Video Scene Graph Generation (VSGG) typically adopt an offline pipeline. Though demonstrating promising performance, they remain unable to handle real-tim…

cs.CV2024

Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models

Liulei Li, Wenguan Wang, Yi Yang

Prevalent human-object interaction (HOI) detection approaches typically leverage large-scale visual-linguistic models to help recognize events involving humans and objects. Though…

cs.CV2022

Deep Hierarchical Semantic Segmentation

Liulei Li, Tianfei Zhou, Wenguan Wang +2

Humans are able to recognize structured relations in observation, allowing us to decompose complex scenes into simpler parts and abstract the visual world in multiple levels. Howev…

cs.CV20222 cited

Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence Learning

Liulei Li, Tianfei Zhou, Wenguan Wang +3

Our target is to learn visual correspondence from unlabeled videos. We develop LIIR, a locality-aware inter-and intra-video reconstruction framework that fills in three missing pie…