activity
20232026
most citedScalable Visual State Space Model with Fractal Scanning

5 citations · 9 across the 31 of their papers we have counts for

collaborators
Showing 2026 · cs.CVShow all

9 papers · 2 filters

cs.CV2026

MotionJEPA: Preventing Temporal Feature Collapse by Capturing Visual Changes in Latent Space

Markus Karmann, Shile Li, Christian Internò +10

Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction. However, standard JEPA train…

cs.CV2026

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Zitong Xu, Huiyu Duan, Xinyun Zhang +7

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstr…

cs.CV2026

Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework

Linxiao Shi, Siming Zheng, Zerong Wang +5

Existing mobile devices are constrained by compact optical designs, such as small apertures, which make it difficult to produce natural, optically realistic bokeh effects. Although…

cs.CV2026

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

Zhiheng Li, Zongyang Ma, Jiaxian Chen +12

The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose t…

cs.CV2026

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

Zitong Xu, Huiyu Duan, Yifei Nie +9

Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as unnatural objects, lighting…

cs.CV2026

Visual Text Compression as Measure Transport

Lv Tang, Tianyi Zheng, Yang Liu +2

Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing --$20\t…