2 citations · 4 across the 3 of their papers we have counts for
3 papers
Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
Théo Gigant, Camille Guinaudeau, Frédéric Dufaux
Vision-Language Models (VLMs) can process visual and textual information in multiple formats: texts, images, interleaved texts and images, or even hour-long videos. In this work, w…
DA-Flow: Dual Attention Normalizing Flow for Skeleton-based Video Anomaly Detection
Ruituo Wu, Yang Chen, Jian Xiao +5
Cooperation between temporal convolutional networks (TCN) and graph convolutional networks (GCN) as a processing module has shown promising results in skeleton-based video anomaly…
Quality evaluation of point clouds: a novel no-reference approach using transformer-based architecture
Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise +1
With the increased interest in immersive experiences, point cloud came to birth and was widely adopted as the first choice to represent 3D media. Besides several distortions that c…