3 papers
cs.CV2026
Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation
Simone Giano, Lorenzo Severini, Alessandro Galdelli +1
The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for transferring knowledge via Da…
cs.CV2026
Structural Kolmogorov-Arnold Convolutions: Learnable Function on the Values or the Filter Shape as Parameter-Efficient Alternative to Per-Edge Convolutional KANs
Stefano Mereu, Oleksandr Kuznetsov, Gabriele Marchello +4
Convolutional Kolmogorov--Arnold Networks (KANs) replace the fixed weights of a convolutional kernel with learnable univariate functions. The dominant formulation attaches one such…
cs.HC2025
Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations
Lorenzo Stacchio, Andrea Ubaldi, Alessandro Galdelli +3
We present Empathic Prompting, a novel framework for multimodal human-AI interaction that enriches Large Language Model (LLM) conversations with implicit non-verbal context. The sy…