activity
20242026
most citedEgoSpot:Egocentric Multimodal Control for Hands-Free Mobile Manipulation

1 citations · 2 across the 31 of their papers we have counts for

collaborators
Showing cs.CVShow all

81 papers · 1 filter

cs.CV2026

X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis

Sonali Godavarthy, Matthias Neuwirth-Trapp, Tim-Felix Faasch +4

Imaging factor disentanglement in text-to-image generation aims to independently control image acquisition properties such as types of camera lenses, sensor types, viewpoints, and…

cs.CV2026

Event-Based Motion Estimation via Oriented Distance Fields

Lei Sun, Yuqin Ma, Weilun Li +5

Event-based motion estimation is central to tasks that demand high temporal resolution and robustness to fast motion. Existing methods typically rely on iterative optimization or r…

cs.CV2026

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Yuqian Fu, Tianwen Qian, Yanjun Li +30

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…

cs.CV2026

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2

Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this are…

cs.CV2026

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

Mohammad Mahdi, Nedko Savov, Danda Pani Paudel +1

Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, syn…

cs.CV2026

InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing

Yebin Yang, Di Wen, Lei Qi +10

Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity…