4 papers
SceneVGGT: VGGT-based online 3D semantic SLAM for indoor scene understanding and navigation
Anna Gelencsér-Horváth, Gergely Dinya, Dorka Boglárka ErÅs +3
We present SceneVGGT, a spatio-temporal 3D scene understanding framework that combines SLAM with semantic mapping for autonomous and assistive navigation. Built on VGGT, our method…
SAMannot: A Memory-Efficient, Local, Open-source Framework for Interactive Video Instance Segmentation based on SAM2
Gergely Dinya, András Gelencsér, Krisztina Kupán +3
Current research workflows for precise video segmentation are often forced into a compromise between labor-intensive manual curation, costly commercial platforms, and/or privacy-co…
Building temporally coherent 3D maps with VGGT for memory-efficient Semantic SLAM
Gergely Dinya, Péter Halász, András LÅrincz +2
We present a fast, spatio-temporal scene understanding framework based on Visual Geometry Grounded Transformer (VGGT). The proposed pipeline is designed to enable efficient, close…
Automatic camera orientation estimation for a partially calibrated camera above a plane with a line at known planar distance
Gergely Dinya, Anna Gelencsér-Horváth
We present a derivation for estimating the roll and pitch orientation of a partially calibrated camera mounted above a planar surface, using minimal scene information. Specifically…