6 papers
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
Le Jiang, Xiangyu Bai, Bishoy Galoaa +7
We present PanoWorld, a panoramic video world model that generates geometry-consistent 360 video from a single image and a caption. Existing panoramic video methods optimi…
Motion-o: Trajectory-Grounded Video Reasoning
Bishoy Galoaa, Shayda Moezzi, Xiangyu Bai +1
Recent video reasoning models increasingly produce spatio-temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grou…
Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
Shayda Moezzi, Bishoy Galoaa, Lorena Genua +2
Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Ju…
Broadening View Synthesis of Dynamic Scenes from Constrained Monocular Videos
Le Jiang, Shaotong Zhu, Yedi Luo +2
In dynamic Neural Radiance Fields (NeRF) systems, state-of-the-art novel view synthesis methods often fail under significant viewpoint deviations, producing unstable and unrealisti…
MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
Xiangyu Bai, He Liang, Bishoy Galoaa +4
While text-to-video (T2V) generation has achieved remarkable progress in photorealism, generating intent-aligned videos that faithfully obey physics principles remains a core chall…
Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers
Bishoy Galoaa, Xiangyu Bai, Shayda Moezzi +4
This paper presents LAPA (Look Around and Pay Attention), a novel end-to-end transformer-based architecture for multi-camera point tracking that integrates appearance-based matchin…