1 citations · 2 across the 31 of their papers we have counts for
81 papers · 1 filter
X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis
Sonali Godavarthy, Matthias Neuwirth-Trapp, Tim-Felix Faasch +4
Imaging factor disentanglement in text-to-image generation aims to independently control image acquisition properties such as types of camera lenses, sensor types, viewpoints, and…
Event-Based Motion Estimation via Oriented Distance Fields
Lei Sun, Yuqin Ma, Weilun Li +5
Event-based motion estimation is central to tasks that demand high temporal resolution and robustness to fast motion. Existing methods typically rely on iterative optimization or r…
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
Yuqian Fu, Tianwen Qian, Yanjun Li +30
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this are…
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
Mohammad Mahdi, Nedko Savov, Danda Pani Paudel +1
Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, syn…
InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing
Yebin Yang, Di Wen, Lei Qi +10
Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity…