8 citations · 8 across the 2 of their papers we have counts for
3 papers
cs.CV2020
Top-1 CORSMAL Challenge 2020 Submission: Filling Mass Estimation Using Multi-modal Observations of Human-robot Handovers
Vladimir Iashin, Francesca Palermo, Gökhan Solak +1
Human-robot object handover is a key skill for the future of human-robot collaboration. CORSMAL 2020 Challenge focuses on the perception part of this problem: the robot needs to es…
cs.CV2020★ 8 cited
A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
Vladimir Iashin, Esa Rahtu
Dense video captioning aims to localize and describe important events in untrimmed videos. Existing methods mainly tackle this task by exploiting only visual features, while comple…
cs.CV2020
Multi-modal Dense Video Captioning
Vladimir Iashin, Esa Rahtu
Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previou…