30 citations · 213 across the 25 of their papers we have counts for
3 papers · 2 filters
DualFormer: Local-Global Stratified Transformer for Efficient Video Recognition
Yuxuan Liang, Pan Zhou, Roger Zimmermann +1
While transformers have shown great potential on video recognition with their strong capability of capturing long-range dependencies, they often suffer high computational costs ind…
Recovering the Unbiased Scene Graphs from the Biased Ones
Meng-Jiun Chiou, Henghui Ding, Hanshu Yan +3
Given input images, scene graph generation (SGG) aims to produce comprehensive, graphical representations describing visual relationships among salient objects. Recently, more effo…
ST-HOI: A Spatial-Temporal Baseline for Human-Object Interaction Detection in Videos
Meng-Jiun Chiou, Chun-Yu Liao, Li-Wei Wang +2
Detecting human-object interactions (HOI) is an important step toward a comprehensive visual understanding of machines. While detecting non-temporal HOIs (e.g., sitting on a chair)…