11 papers
Sense it with your eyes: Sensation Generation and Understanding for Advertisements
Aysan Aghazadeh, Sina Malakouti, Adriana Kovashka
Sensory advertising evokes human senses through visual cues, enabling audiences to mentally simulate experiences and increasing persuasive impact. Despite the recent increase in us…
Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching
Arushi Rai, Adriana Kovashka
Video-LLMs often attend to irrelevant frames, which is especially detrimental for sports coaching tasks requiring precise temporal grounding. Yet obtaining frame-level supervision…
Culture in Action: Evaluating Text-to-Image Models through Social Activities
Sina Malakouti, Boqing Gong, Adriana Kovashka
Text-to-image (T2I) diffusion models achieve impressive photorealism by training on large-scale web data, but models inherit cultural biases and fail to depict underrepresented reg…
Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case Study
Arushi Rai, Adriana Kovashka
While there is rapid progress in video-LLMs with advanced reasoning capabilities, prior work shows that these models struggle on the challenging task of sports feedback generation…
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
Kyle Buettner, Jacob T. Emmerson, Adriana Kovashka
When captioning an image, people describe objects in diverse ways, such as by using different terms and/or including details that are perceptually noteworthy to them. Descriptions…
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition
Sina Malakouti, Adriana Kovashka
Text-to-image (T2I) diffusion models exhibit impressive photorealistic image generation capabilities, yet they struggle in compositional image generation. In this work, we introduc…