activity
20172026
most citedReal-Time Highly Accurate Dense Depth on a Power Budget using an FPGA-CPU Hybrid SoC

16 citations · 28 across the 23 of their papers we have counts for

collaborators
Showing cs.CVShow all

26 papers · 1 filter

cs.CV2026

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Yuqian Fu, Tianwen Qian, Yanjun Li +30

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…

cs.CV2026

Editing Everything Everywhere All at Once

Fabio Quattrini, Carmine Zaccagnino, Enis Simsar +4

Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency and potentially better harm…

cs.CV2026

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr +2

Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal large language models (MLLMs). Ho…

cs.CV2026

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr +2

Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they still struggle with object hall…

cs.CV2026

MOSAIC-GS: Monocular Scene Reconstruction via Advanced Initialization for Complex Dynamic Environments

Svitlana Morkva, Maximum Wilder-Smith, Michael Oechsle +3

We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos using Gaussian Splatting.…

cs.CV2025

Autoregressive Styled Text Image Generation, but Make it Reliable

Carmine Zaccagnino, Fabio Quattrini, Vittorio Pippi +3

Generating faithful and readable styled text images (especially for Styled Handwritten Text generation - HTG) is an open problem with several possible applications across graphic d…