7 papers
Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories
Yixuan Yang, Mehak Arora, Ryan Zhang +10
We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories. JEPA architectures have enabled latent-spac…
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
Huichan Seo, Sieun Choi, Minki Hong +8
Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I) systems, leaving image-to-image…
RIO: Flexible Real-Time Robot I/O for Cross-Embodiment Robot Learning
Pablo Ortega-Kral, Eliot Xing, Arthur Bucker +13
Despite recent efforts to collect multi-task, multi-embodiment datasets, to design recipes for training Vision-Language-Action models (VLAs), and to showcase these models on differ…
MobileOcc: A Human-Aware Semantic Occupancy Dataset for Mobile Robots
Junseo Kim, Guido Dumont, Xinyu Gao +3
Dense 3D semantic occupancy perception is critical for mobile robots operating in pedestrian-rich environments, yet it remains underexplored compared to its application in autonomo…
PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings
Junseo Kim, Jongwook Han, Dongmin Choi +3
Visual persuasion, which uses visual elements to influence cognition and behaviors, is crucial in fields such as advertising and political communication. With recent advancements i…
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
Jaewoo Ahn, Junseo Kim, Heeseung Yun +4
GUI agents powered by LLMs show promise in interacting with diverse digital environments. Among these, video games offer a valuable testbed due to their varied interfaces, with adv…