3 papers
cs.CV2026
Multi-modal video data-pipelines for machine learning with minimal human supervision
Mihai-Cristian Pîrvu, Marius Leordeanu
The real-world is inherently multi-modal at its core. Our tools observe and take snapshots of it, in digital form, such as videos or sounds, however much of it is lost. Similarly f…
cs.CV2025
With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
Ciprian Constantinescu, Marius Leordeanu
Humans effortlessly identify objects by leveraging a rich understanding of the surrounding scene, including spatial relationships, material properties, and the co-occurrence of oth…
cs.CV2025
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
Pîrvu Mihai-Cristian, Marius Leordeanu
The computer vision domain has greatly benefited from an abundance of data across many modalities to improve on various visual tasks. Recently, there has been a lot of focus on sel…