14 papers
GeoTrace: Geometry-Aware Trajectory Token Compression for Video Large Language Models
Guohuan Xie, Mengqi Lei, Chuan Shi +3
Although Video Large Language Models (Video LLMs) have shown strong performance in video understanding, their efficiency is still limited by the large number of visual tokens. Exis…
H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
Yongji Zhang, Siqi Li, Kuiyang Huang +2
Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations. Existing approaches typically rely on f…
M^2C-EvDet: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection
Wei Bao, Siqi Li, Shouan Pan +2
Event-based object Detection (EvDet), as a biologically inspired visual perception paradigm, demonstrates superior performance in scenarios demanding high temporal resolution and a…
Count Anything
Mengqi Lei, Shuokun Cheng, Wei Bao +4
Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tai…
Hypergraph as Language
Mengqi Lei, Guohuan Xie, Shihui Ying +6
Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally graph-centric: they focus on…
Rethinking Event-Based Object Detection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning
Meisen Wang, Hao Deng, Wei Bao +6
Event cameras provide microsecond-level temporal resolution, low latency, and high dynamic range, offering potential for perception under fast motion and challenging illumination c…