coarse-to-fine modeling 1motion modeling 1spatio-temporal video grounding 1temporal localization 1vision-language fusion 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding
Kai Chen, Ming Dai, Wenxuan Cheng +1
The paper introduces ScanFocus, a coarse-to-fine framework for spatio-temporal video grounding that first scans videos globally to generate coarse object proposals and then refines…
cs.AI2026
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
Xianhui Meng, Zirui Song, Yuchen Zhang +10
Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible…