2 papers
cs.CV2025
VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism
Congzhi Zhang, Jiawei Peng, Zhenglin Wang +5
Large Vision-Language Models (LVLMs) have shown exceptional performance in multimodal tasks, but their effectiveness in complex visual reasoning is still constrained, especially wh…
cs.CV2025
MuseFace: Text-driven Face Editing via Diffusion-based Mask Generation Approach
Xin Zhang, Siting Huang, Xiangyang Luo +5
Face editing modifies the appearance of face, which plays a key role in customization and enhancement of personal images. Although much work have achieved remarkable success in tex…