3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2024
Towards Event-oriented Long Video Understanding
Yifan Du, Kun Zhou, Yuqi Huo +7
With the rapid development of video Multimodal Large Language Models (MLLMs), numerous benchmarks have been proposed to assess their video understanding capability. However, due to…
cs.CV2024
Temporal Adaptive RGBT Tracking with Modality Prompt
Hongyu Wang, Xiaotao Liu, Yifan Li +3
RGBT tracking has been widely used in various fields such as robotics, surveillance processing, and autonomous driving. Existing RGBT trackers fully explore the spatial information…
cs.RO2023★ 3 cited
WALL-E: Embodied Robotic WAiter Load Lifting with Large Language Model
Tianyu Wang, Yifan Li, Haitao Lin +2
Enabling robots to understand language instructions and react accordingly to visual perception has been a long-standing goal in the robotics research community. Achieving this goal…