6 papers
Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs
Agnese Taluzzi, Riccardo Santambrogio, Simone Mentasti +2
Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations. As a result, current video…
SGR-OCC: Evolving Monocular Priors for Embodied 3D Occupancy Prediction via Soft-Gating Lifting and Semantic-Adaptive Geometric Refinement
Yiran Guo, Simone Mentasti, Xiaofeng Jin +2
3D semantic occupancy prediction is a cornerstone for embodied AI, enabling agents to perceive dense scene geometry and semantics incrementally from monocular video streams. Howeve…
EETnet: a CNN for Gaze Detection and Tracking for Smart-Eyewear
Andrea Aspesi, Andrea Simpsi, Aaron Tognoli +3
Event-based cameras are becoming a popular solution for efficient, low-power eye tracking. Due to the sparse and asynchronous nature of event data, they require less processing pow…
From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge
Agnese Taluzzi, Davide Gesualdi, Riccardo Santambrogio +4
This report presents SceneNet and KnowledgeNet, our approaches developed for the HD-EPIC VQA Challenge 2025. SceneNet leverages scene graphs generated with a multi-modal large lang…
High-frequency near-eye ground truth for event-based eye tracking
Andrea Simpsi, Andrea Aspesi, Simone Mentasti +3
Event-based eye tracking is a promising solution for efficient and low-power eye tracking in smart eyewear technologies. However, the novelty of event-based sensors has resulted in…
A Spatio-temporal Graph Network Allowing Incomplete Trajectory Input for Pedestrian Trajectory Prediction
Juncen Long, Gianluca Bardaro, Simone Mentasti +1
Pedestrian trajectory prediction is important in the research of mobile robot navigation in environments with pedestrians. Most pedestrian trajectory prediction algorithms require…