activity
20242026
collaborators

101 papers

cs.CV2026

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Yuqian Fu, Tianwen Qian, Yanjun Li +30

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…

cs.CV2026

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2

Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this are…

cs.AI2026

Knowledge-Centric Agents for Workflow Generation in ComfyUI

Zhendong Li, Lei Sun, Ruibo Ming +4

Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large languag…

eess.IV2026

Towards Blind Lens Aberration Correction via Large LensLib Pre-training and Discrete Degradation Priors

Xiaolong Qian, Qi Jiang, Yao Gao +8

Emerging deep-learning-based lens library pre-training (LensLib-PT) pipeline offers a new avenue for blind lens aberration correction by training a universal neural network, demons…

cs.CV2026

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

Mohammad Mahdi, Nedko Savov, Danda Pani Paudel +1

Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, syn…

cs.CV2026

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation

Edoardo Mello Rella, Ajad Chhatkuli, Shipra Jain +2

Recent online video instance segmentation (VIS) methods have achieved impressive results, thus becoming the preferred approach to segment instances in videos. Despite the resurgenc…