5 papers
UniTrack: Differentiable Graph Representation Learning for Multi-Object Tracking
Bishoy Galoaa, Xiangyu Bai, Utsav Nandi +3
We present UniTrack, a plug-and-play graph-theoretic loss function designed to significantly enhance multi-object tracking (MOT) performance by directly optimizing tracking-specifi…
Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces
Bishoy Galoaa, Xiangyu Bai, Sarah Ostadabbas
We present Lang2Motion, a framework for language-guided point trajectory generation by aligning motion manifolds with joint embedding spaces. Unlike prior work focusing on human mo…
MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
Xiangyu Bai, He Liang, Bishoy Galoaa +4
While text-to-video (T2V) generation has achieved remarkable progress in photorealism, generating intent-aligned videos that faithfully obey physics principles remains a core chall…
Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers
Bishoy Galoaa, Xiangyu Bai, Shayda Moezzi +4
This paper presents LAPA (Look Around and Pay Attention), a novel end-to-end transformer-based architecture for multi-camera point tracking that integrates appearance-based matchin…
Learning Multimodal AI Algorithms for Amplifying Limited User Input into High-dimensional Control Space
Ali Rabiee, Sima Ghafoori, MH Farhadi +5
Current invasive assistive technologies are designed to infer high-dimensional motor control signals from severely paralyzed patients. However, they face significant challenges, in…