most citedGSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement

40 citations · 79 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI2025

HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions

Yifei Dong, Fengyi Wu, Qi He +9

Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, crowded environments. We present HA-VLN 2.0,…

cs.CV2022★ 31 cited

Hypergraph Transformer for Skeleton-based Action Recognition

Yuxuan Zhou, Zhi-Qi Cheng, Chao Li +4

Skeleton-based action recognition aims to recognize human actions given human joint coordinates with skeletal interconnections. By defining a graph with joints as vertices and thei…

cs.CV2022★ 2 cited

LongShortNet: Exploring Temporal and Semantic Features Fusion in Streaming Perception

Chenyang Li, Zhi-Qi Cheng, Jun-Yan He +6

Streaming perception is a critical task in autonomous driving that requires balancing the latency and accuracy of the autopilot system. However, current methods for streaming perce…

cs.CV2022★ 6 cited

ProContEXT: Exploring Progressive Context Transformer for Tracking

Jin-Peng Lan, Zhi-Qi Cheng, Jun-Yan He +6

Existing Visual Object Tracking (VOT) only takes the target area in the first frame as a template. This causes tracking to inevitably fail in fast-changing and crowded scenes, as i…

cs.CV2022★ 40 cited

GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement

Zhi-Qi Cheng, Qi Dai, Siyao Li +2

Grounded Situation Recognition (GSR) aims to generate structured semantic summaries of images for "human-like" event understanding. Specifically, GSR task not only detects the sali…