collaborators

6 papers

cs.CV2026

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs

Agnese Taluzzi, Riccardo Santambrogio, Simone Mentasti +2

Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations. As a result, current video…

cs.CV2026

SGR-OCC: Evolving Monocular Priors for Embodied 3D Occupancy Prediction via Soft-Gating Lifting and Semantic-Adaptive Geometric Refinement

Yiran Guo, Simone Mentasti, Xiaofeng Jin +2

3D semantic occupancy prediction is a cornerstone for embodied AI, enabling agents to perceive dense scene geometry and semantics incrementally from monocular video streams. Howeve…

cs.CV2025

EETnet: a CNN for Gaze Detection and Tracking for Smart-Eyewear

Andrea Aspesi, Andrea Simpsi, Aaron Tognoli +3

Event-based cameras are becoming a popular solution for efficient, low-power eye tracking. Due to the sparse and asynchronous nature of event data, they require less processing pow…

cs.CV2025

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

Agnese Taluzzi, Davide Gesualdi, Riccardo Santambrogio +4

This report presents SceneNet and KnowledgeNet, our approaches developed for the HD-EPIC VQA Challenge 2025. SceneNet leverages scene graphs generated with a multi-modal large lang…

cs.CV2025

High-frequency near-eye ground truth for event-based eye tracking

Andrea Simpsi, Andrea Aspesi, Simone Mentasti +3

Event-based eye tracking is a promising solution for efficient and low-power eye tracking in smart eyewear technologies. However, the novelty of event-based sensors has resulted in…

cs.CV2025

A Spatio-temporal Graph Network Allowing Incomplete Trajectory Input for Pedestrian Trajectory Prediction

Juncen Long, Gianluca Bardaro, Simone Mentasti +1

Pedestrian trajectory prediction is important in the research of mobile robot navigation in environments with pedestrians. Most pedestrian trajectory prediction algorithms require…