13 papers · 1 filter
MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs
Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche +7
Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: how fast something moves, whic…
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
Le Jiang, Xiangyu Bai, Bishoy Galoaa +7
We present PanoWorld, a panoramic video world model that generates geometry-consistent 360 video from a single image and a caption. Existing panoramic video methods optimi…
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Juyi Lin, Arash Akbari, Yumei He +13
Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, ev…
HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models
Xiangyu Bai, Bishoy Galoaa, Sarah Ostadabbas
Video question answering (VQA) with vision-language models (VLMs) depends critically on which frames are selected from the input video, yet most systems rely on uniform or heuristi…
Motion-o: Trajectory-Grounded Video Reasoning
Bishoy Galoaa, Shayda Moezzi, Xiangyu Bai +1
Recent video reasoning models increasingly produce spatio-temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grou…
UniTrack: Differentiable Graph Representation Learning for Multi-Object Tracking
Bishoy Galoaa, Xiangyu Bai, Utsav Nandi +3
We present UniTrack, a plug-and-play graph-theoretic loss function designed to significantly enhance multi-object tracking (MOT) performance by directly optimizing tracking-specifi…