4 papers
What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability
Lucia Cascone, Valeria Fraenza, Michele Nappi +2
Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input…
Bias Analysis for Synthetic Face Detection: A Case Study of the Impact of Facial Attributes
Asmae Lamsaf, Lucia Cascone, Hugo Proença +1
Bias analysis for synthetic face detection is bound to become a critical topic in the coming years. Although many detection models have been developed and several datasets have bee…
Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism
Chang Zong, Hang Zhou
The accurate prediction of stock movements is crucial for investment strategies. Stock prices are subject to the influence of various forms of information, including financial indi…
Ollivier-Ricci Curvature For Head Pose Estimation From a Single Image
Lucia Cascone, Riccardo Distasi, Michele Nappi
Head pose estimation is a crucial challenge for many real-world applications, such as attention and human behavior analysis. This paper aims to estimate head pose from a single ima…