1 paper
Pengtao Ma, Ziliang Zhou, Ciyu Ruan +7
First-person dynamic spatial reasoning requires models to track continuous motion and precise geometric structure, but the quadratic attention cost of Transformer-based Video-LLMs…