2 papers
cs.CV2024
Caption-Driven Explorations: Aligning Image and Text Embeddings through Human-Inspired Foveated Vision
Dario Zanca, Andrea Zugarini, Simon Dietz +4
Understanding human attention is crucial for vision science and AI. While many models exist for free-viewing, less is known about task-driven image exploration. To address this, we…
cs.LG2024
How Intermodal Interaction Affects the Performance of Deep Multimodal Fusion for Mixed-Type Time Series
Simon Dietz, Thomas Altstidl, Dario Zanca +2
Mixed-type time series (MTTS) is a bimodal data type that is common in many domains, such as healthcare, finance, environmental monitoring, and social media. It consists of regular…