1 paper · 1 filter
Ee Yeo Keat, Zhang Hao, Alexander Matyasko +1
We introduce VidTFS, a Training-free, open-vocabulary video goal and action inference framework that combines the frozen vision foundational model (VFM) and large language model (L…