5 papers
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
Qihuang Zhong, Liang Ding, Wenjie Xuan +3
Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acquiring high-quality reasoning…
Detect Changes like Humans: Incorporating Semantic Priors for Improved Change Detection
Yuhang Gan, Wenjie Xuan, Zhiming Luo +4
When given two similar images, humans identify their differences by comparing the appearance (e.g., color, texture) with the help of semantics (e.g., objects, relations). However,…
Rethink Sparse Signals for Pose-guided Text-to-image Generation
Wenjie Xuan, Jing Zhang, Juhua Liu +2
Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-imag…
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
Wenjie Xuan, Yufei Xu, Shanshan Zhao +4
ControlNet excels at creating content that closely matches precise contours in user-provided masks. However, when these masks contain noise, as a frequent occurrence with non-exper…
RFL-CDNet: Towards Accurate Change Detection via Richer Feature Learning
Yuhang Gan, Wenjie Xuan, Hang Chen +2
Change Detection is a crucial but extremely challenging task of remote sensing image analysis, and much progress has been made with the rapid development of deep learning. However,…