3 papers
cs.CV2026
Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos
Liu Liu, Freya Huying Tan, Fábio Duarte
We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,0…
cs.CV2026
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
Liu Liu, Alexandra Kudaeva, Marco Cipriano +4
Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such in…
cs.CV2024
ELSA: Evaluating Localization of Social Activities in Urban Streets using Open-Vocabulary Detection
Maryam Hosseini, Marco Cipriano, Sedigheh Eslami +4
Existing Open Vocabulary Detection (OVD) models exhibit a number of challenges. They often struggle with semantic consistency across diverse inputs, and are often sensitive to slig…