6 papers
MDiff4STR: Mask Diffusion Model for Scene Text Recognition
Yongkun Du, Miaomiao Zhao, Songlin Fan +3
Mask Diffusion Models (MDMs) have recently emerged as a promising alternative to auto-regressive models (ARMs) for vision-language tasks, owing to their flexible balance of efficie…
VGD: Visual Geometry Gaussian Splatting for Feed-Forward Surround-view Driving Reconstruction
Junhong Lin, Kangli Wang, Shunzhou Wang +3
Feed-forward surround-view autonomous driving scene reconstruction offers fast, generalizable inference ability, which faces the core challenge of ensuring generalization while ele…
Stochasticity-aware No-Reference Point Cloud Quality Assessment
Songlin Fan, Wei Gao, Zhineng Chen +3
The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) a…
Consistent Video Editing as Flow-Driven Image-to-Video Generation
Ge Wang, Songlin Fan, Hangxu Liu +3
With the prosper of video diffusion models, down-stream applications like video editing have been significantly promoted without consuming much computational cost. One particular c…
IE-Bench: Advancing the Measurement of Text-Driven Image Editing for Human Perception Alignment
Shangkun Sun, Bowen Qu, Xiaoyu Liang +2
Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different…
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
Shangkun Sun, Xiaoyu Liang, Songlin Fan +2
Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align…