3 citations · 4 across the 6 of their papers we have counts for
7 papers · 1 filter
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
Bin Li, Shenxi Liu, Yixuan Weng +3
Following the successful hosts of the 1-st (NLPCC 2023 Foshan) CMIVQA and the 2-rd (NLPCC 2024 Hangzhou) MMIVQA challenges, this year, a new task has been introduced to further adv…
PICD: Versatile Perceptual Image Compression with Diffusion Rendering
Tongda Xu, Jiahao Li, Bin Li +3
Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existi…
Rethinking Person Re-identification from a Projection-on-Prototypes Perspective
Qizao Wang, Xuelin Qian, Bin Li +2
Person Re-IDentification (Re-ID) as a retrieval task, has achieved tremendous development over the past decade. Existing state-of-the-art methods follow an analogous framework to f…
DeFeeNet: Consecutive 3D Human Motion Prediction with Deviation Feedback
Xiaoning Sun, Huaijiang Sun, Bin Li +3
Let us rethink the real-world scenarios that require human motion prediction techniques, such as human-robot collaboration. Current works simplify the task of predicting human moti…
Meta-Auxiliary Learning for Adaptive Human Pose Prediction
Qiongjie Cui, Huaijiang Sun, Jianfeng Lu +2
Predicting high-fidelity future human poses, from a historically observed sequence, is decisive for intelligent robots to interact with humans. Deep end-to-end learning approaches,…
Weakly-Supervised Text Instance Segmentation
Xinyan Zu, Haiyang Yu, Bin Li +1
Text segmentation is a challenging vision task with many downstream applications. Current text segmentation methods require pixel-level annotations, which are expensive in the cost…