papers

Publications (8)

cs.CV2024

Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition

Jiaming Zhou, Hanjun Li, Kun-Yu Lin +1

Developing end-to-end action recognition models on long videos is fundamental and crucial for long-video action understanding. Due to the unaffordable cost of end-to-end training o…

cs.CV2022

SIOD: Single Instance Annotated Per Category Per Image for Object Detection

Hanjun Li, Xingjia Pan, Ke Yan +2

Object detection under imperfect data receives great attention recently. Weakly supervised object detection (WSOD) suffers from severe localization issues due to the lack of instan…

cs.CV2023

D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance Annotation

Hanjun Li, Xiujun Shu, Sunan He +5

Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query. Recently, weakly supervised methods still have a lar…

cs.CV2021

Combined Depth Space based Architecture Search For Person Re-identification

Hanjun Li, Gaojie Wu, Wei-Shi Zheng

Most works on person re-identification (ReID) take advantage of large backbone networks such as ResNet, which are designed for image classification instead of ReID, for feature ext…

cs.CV2024

Unified and Dynamic Graph for Temporal Character Grouping in Long Videos

Xiujun Shu, Wei Wen, Liangsheng Xu +6

Video temporal character grouping locates appearing moments of major characters within a video according to their identities. To this end, recent works have evolved from unsupervis…

cs.CV2023

Collaborative Noisy Label Cleaner: Learning Scene-aware Trailers for Multi-modal Highlight Detection in Movies

Bei Gan, Xiujun Shu, Ruizhi Qiao +4

Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1…