activity
20242026
collaborators

84 papers

cs.AI2026

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel +2

Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably…

cs.CV2026

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Yuqian Fu, Tianwen Qian, Yanjun Li +30

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…

cs.CV2026

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2

Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this are…

cs.AI2026

Knowledge-Centric Agents for Workflow Generation in ComfyUI

Zhendong Li, Lei Sun, Ruibo Ming +4

Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large languag…

cs.CV2026

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

Mohammad Mahdi, Nedko Savov, Danda Pani Paudel +1

Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, syn…

cs.CV2026

InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing

Yebin Yang, Di Wen, Lei Qi +10

Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity…