works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.CV2026

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation

Bishoy Galoaa, Sarah Ostadabbas

The paper introduces TCAM, a system that automatically detects moving objects in videos, tracks their precise point trajectories, and generates open‑vocabulary captions describing…

cs.RO2026

Uncertainty-Aware Ankle Exoskeleton Control

Fatima Mumtaza Tourk, Bishoy Galoaa, Sanat Shajan +3

Lower limb exoskeletons show promise to assist human movement, but their utility is limited by controllers designed for discrete, predefined actions in controlled environments, res…

cs.CV2026

PanoWorld: Geometry-Consistent Panoramic Video World Modeling

Le Jiang, Xiangyu Bai, Bishoy Galoaa +7

We present PanoWorld, a panoramic video world model that generates geometry-consistent 360 video from a single image and a caption. Existing panoramic video methods optimi…

cs.CV2026

Structure Over Scale: Learning Visual Reasoning from Pedagogical Video

Bishoy Galoaa, Xiangyu Bai, Sarah Ostadabbas

State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial relations, navigation, and objec…

cs.CV2026

Motion-o: Trajectory-Grounded Video Reasoning

Bishoy Galoaa, Shayda Moezzi, Xiangyu Bai +1

Recent video reasoning models increasingly produce spatio-temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grou…

cs.CL2026

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

Shayda Moezzi, Bishoy Galoaa, Lorena Genua +2

Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Ju…