3 papers
cs.CV2026
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
Qihuang Zhong, Liang Ding, Wenjie Xuan +3
Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acquiring high-quality reasoning…
cs.CV2025
Detect Changes like Humans: Incorporating Semantic Priors for Improved Change Detection
Yuhang Gan, Wenjie Xuan, Zhiming Luo +4
When given two similar images, humans identify their differences by comparing the appearance (e.g., color, texture) with the help of semantics (e.g., objects, relations). However,…
cs.CV2025
Rethink Sparse Signals for Pose-guided Text-to-image Generation
Wenjie Xuan, Jing Zhang, Juhua Liu +2
Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-imag…