1 paper
Sang Min Kim, Hyeongjun Heo, Junho Kim +2
We propose Point2Act, which directly retrieves the 3D action point relevant to a contextually described task, leveraging Multimodal Large Language Models (MLLMs). Foundation models…