most citedSemantic Correspondence: Unified Benchmarking and a Strong Baseline

2 citations · 2 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV2026

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma +31

Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they a…

cs.CV2026

Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation

Jingyi Lu, Kai Han

Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpaint…

cs.CV20262 cited

Semantic Correspondence: Unified Benchmarking and a Strong Baseline

Kaiyan Zhang, Xinghui Li, Jingyi Lu +1

Establishing semantic correspondence is a challenging task in computer vision, aiming to match keypoints with the same semantic information across different images. Benefiting from…

cs.CV2026

Semantic-Enriched Latent Visual Reasoning

Tianrun Xu, Yue Sun, Qixun Wang +8

Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…

cs.CV2026

VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation

Yan Luo, Ahmadou Aidara, Jingyi Lu +3

Classifier-free guidance (CFG) is the primary control over how strongly text semantics move a flow-based sampler, yet standard practice holds its scale fixed across the entire ODE…

cs.CV2025

Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping

Jingyi Lu, Kai Han

Drag-based image editing has emerged as a powerful paradigm for intuitive image manipulation. However, existing approaches predominantly rely on manipulating the latent space of ge…