4 papers
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs
Lifan Jiang, Tianrun Wu, Yuhang Pei +3
The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms, making fair cross-paradigm c…
GeoSeg: Training-Free Reasoning-Driven Segmentation in Remote Sensing Imagery
Lifan Jiang, Yuhang Pei, oxi Wu +5
Recent advances in MLLMs are reframing segmentation from fixed-category prediction to instruction-grounded localization. While reasoning based segmentation has progressed rapidly i…
RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization
Linxuan Xia, Xiaolong Yang, Yongyuan Chen +4
Aligning large language models (LLMs) on domain-specific data remains a fundamental challenge. Supervised fine-tuning (SFT) offers a straightforward way to inject domain knowledge…
SNR-Edit: Structure-Aware Noise Rectification for Inversion-Free Flow-Based Editing
Lifan Jiang, Boxi Wu, Yuhang Pei +5
Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing approaches rely on fixed Gaussian noise to co…