6 papers
GeoTrace: Geometry-Aware Trajectory Token Compression for Video Large Language Models
Guohuan Xie, Mengqi Lei, Chuan Shi +3
Although Video Large Language Models (Video LLMs) have shown strong performance in video understanding, their efficiency is still limited by the large number of visual tokens. Exis…
Count Anything
Mengqi Lei, Shuokun Cheng, Wei Bao +4
Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tai…
Hypergraph as Language
Mengqi Lei, Guohuan Xie, Shihui Ying +6
Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally graph-centric: they focus on…
RACANet: Reliability-Aware Crowd Anchor Network for RGB-T Crowd Counting
Jinghao Shi, Mengqi Lei, Kunliang He +3
RGB-Thermal (T) crowd counting aims to integrate visible-spectrum and thermal infrared information to improve the robustness of crowd density estimation in complex scenes. Although…
SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
Mengqi Lei, Yihong Wu, Siqi Li +4
Visual recognition relies on understanding the semantics of image tokens and their complex interactions. Mainstream self-attention methods, while effective at modeling global pair-…
YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
Mengqi Lei, Siqi Li, Yihong Wu +7
The YOLO series models reign supreme in real-time object detection due to their superior accuracy and computational efficiency. However, both the convolutional architectures of YOL…