3 papers
cs.CV2025
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
Seokwon Song, Minsu Park, Gunhee Kim
Source attribution aims to enhance the reliability of AI-generated answers by including references for each statement, helping users validate the provided answers. However, existin…
cs.GR2025
Text-Driven Video Style Transfer with State-Space Models: Extending StyleMamba for Temporal Coherence
Chao Li, Minsu Park, Cristina Rossi +1
StyleMamba has recently demonstrated efficient text-driven image style transfer by leveraging state-space models (SSMs) and masked directional losses. In this paper, we extend the…
cs.CY2024
Inclusive content reduces racial and gender biases, yet non-inclusive content dominates popular culture
Nouar AlDahoul, Hazem Ibrahim, Minsu Park +2
Images are often termed as representations of perceived reality. As such, racial and gender biases in popular culture and visual media could play a critical role in shaping people'…