40 citations · 63 across the 5 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 1 cited
All in Tokens: Unifying Output Space of Visual Tasks via Soft Token
Jia Ning, Chen Li, Zheng Zhang +4
Unlike language tasks, where the output space is usually limited to a set of tokens, the output space of visual tasks is more complicated, making it difficult to build a unified vi…
cs.CV2022★ 40 cited
GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement
Zhi-Qi Cheng, Qi Dai, Siyao Li +2
Grounded Situation Recognition (GSR) aims to generate structured semantic summaries of images for "human-like" event understanding. Specifically, GSR task not only detects the sali…