collaborators

7 papers

cs.CV2026

Depth-Guided Video Object Counting in Crowded Scenes

Yuanjing Xu, Xinyan Liu, Weidong Chen +5

Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Exis…

cs.CV2026

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

Yaowu Fan, Jia Wan, Tao Han +3

Counting and tracking dense crowds in large-scale scenes is a highly practical yet challenging problem. Existing methods mostly rely on fixed-camera datasets with limited scene cov…

cs.CV2026

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

Zelin Zheng, Xinyan Liu, Ruixin Li +4

Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often leading to brittle num…

cs.CV2026

Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes

Qi Zhang, Jixuan Chen, Kaiyi Zhang +3

Multi-view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs-based multi-view crowd tracking archit…

cs.CV2026

TowerDataset: A Heterogeneous Benchmark for Transmission Corridor Segmentation with a Global-Local Fusion Framework

Xu Cui, Xinyan Liu, Chen Yang +4

Fine-grained semantic segmentation of transmission-corridor point clouds is fundamental for intelligent power-line inspection. However, current progress is limited by realistic dat…

cs.CV2026

Dense Point-to-Mask Optimization with Reinforced Point Selection for Crowd Instance Segmentation

Hongru Chen, Jiyang Huang, Jia Wan +1

Crowd instance segmentation is a crucial task with a wide range of applications, including surveillance and transportation. Currently, point labels are common in crowd datasets, wh…