collaborators

6 papers

cs.CV2025

MDiff4STR: Mask Diffusion Model for Scene Text Recognition

Yongkun Du, Miaomiao Zhao, Songlin Fan +3

Mask Diffusion Models (MDMs) have recently emerged as a promising alternative to auto-regressive models (ARMs) for vision-language tasks, owing to their flexible balance of efficie…

cs.CV2025

VGD: Visual Geometry Gaussian Splatting for Feed-Forward Surround-view Driving Reconstruction

Junhong Lin, Kangli Wang, Shunzhou Wang +3

Feed-forward surround-view autonomous driving scene reconstruction offers fast, generalizable inference ability, which faces the core challenge of ensuring generalization while ele…

cs.CV2025

Stochasticity-aware No-Reference Point Cloud Quality Assessment

Songlin Fan, Wei Gao, Zhineng Chen +3

The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) a…

cs.CV2025

Consistent Video Editing as Flow-Driven Image-to-Video Generation

Ge Wang, Songlin Fan, Hangxu Liu +3

With the prosper of video diffusion models, down-stream applications like video editing have been significantly promoted without consuming much computational cost. One particular c…

cs.CV2025

IE-Bench: Advancing the Measurement of Text-Driven Image Editing for Human Perception Alignment

Shangkun Sun, Bowen Qu, Xiaoyu Liang +2

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different…

cs.CV2024

VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment

Shangkun Sun, Xiaoyu Liang, Songlin Fan +2

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align…