84 papers
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel +2
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably…
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
Yuqian Fu, Tianwen Qian, Yanjun Li +30
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this are…
Knowledge-Centric Agents for Workflow Generation in ComfyUI
Zhendong Li, Lei Sun, Ruibo Ming +4
Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large languag…
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
Mohammad Mahdi, Nedko Savov, Danda Pani Paudel +1
Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, syn…
InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing
Yebin Yang, Di Wen, Lei Qi +10
Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity…