From the 1 of 16 linked papers with an AI index.
16 papers
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation
Bishoy Galoaa, Sarah Ostadabbas
The paper introduces TCAM, a system that automatically detects moving objects in videos, tracks their precise point trajectories, and generates open‑vocabulary captions describing…
Human Cognition in Machines: A Unified Perspective of World Models
Timothy Rupprecht, Pu Zhao, Amir Taherin +20
This report of world models distinguishes prior works by the cognitive functions they innovate. Many works claim an almost human-like cognitive capability in their world models. To…
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
Le Jiang, Xiangyu Bai, Bishoy Galoaa +7
We present PanoWorld, a panoramic video world model that generates geometry-consistent 360 video from a single image and a caption. Existing panoramic video methods optimi…
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Juyi Lin, Arash Akbari, Yumei He +13
Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, ev…
Structure Over Scale: Learning Visual Reasoning from Pedagogical Video
Bishoy Galoaa, Xiangyu Bai, Sarah Ostadabbas
State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial relations, navigation, and objec…
Motion-o: Trajectory-Grounded Video Reasoning
Bishoy Galoaa, Shayda Moezzi, Xiangyu Bai +1
Recent video reasoning models increasingly produce spatio-temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grou…