4 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.CV2024
Egocentric Vision Language Planning
Zhirui Fang, Ming Yang, Weishuai Zeng +5
We explore leveraging large multi-modal models (LMMs) and text2image models to build a more general embodied agent. LMMs excel in planning long-horizon tasks over symbolic abstract…
cs.AI2024★ 3 cited
A Survey on Game Playing Agents and Large Models: Methods, Applications, and Challenges
Xinrun Xu, Yuxin Wang, Chaoyi Xu +4
The swift evolution of Large-scale Models (LMs), either language-focused or multi-modal, has garnered extensive attention in both academy and industry. But despite the surge in int…
cs.CV2023★ 4 cited
Unsupervised Optical Flow Estimation with Dynamic Timing Representation for Spike Camera
Lujie Xia, Ziluo Ding, Rui Zhao +5
Efficiently selecting an appropriate spike stream data length to extract precise information is the key to the spike vision tasks. To address this issue, we propose a dynamic timin…