11 papers
RFDM: Residual Flow Diffusion Model for Efficient Causal Video Editing
Mohammadreza Salehi, Mehdi Noroozi, Luca Morreale +4
Instructional video editing applies edits to an input video using only text prompts, enabling intuitive natural-language control. Despite rapid progress, most methods still require…
Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
Valentina Lilova, Toyesh Chakravorty, Julian I. Bibo +5
Benchmarking 3D spatial understanding of foundation models is essential for real-world applications such as robotics and autonomous driving. Existing evaluations often rely on down…
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
Mohammadreza Salehi, Shashanka Venkataramanan, Ioana Simion +3
Dense self-supervised learning has shown great promise for learning pixel- and patch-level representations, but extending it to videos remains challenging due to the complexity of…
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
Shashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi +5
We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance…
Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation
Divyanshu Mishra, Mohammadreza Salehi, Pramit Saha +4
Self-supervised learning (SSL) has achieved major advances in natural images and video understanding, but challenges remain in domains like echocardiography (heart ultrasound) due…
Disjoint chorded cycles in a -connected graph
Zai Ping Lu, Shu Dan Xue
A chorded cycle in a graph is a cycle containing an edge of that joins two nonconsecutive vertices of the cycle. In 2010, Gao and Qiao independently proved that a graph of…