3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2025
BaseReward: A Strong Baseline for Multimodal Reward Model
Yi-Fan Zhang, Haihua Yang, Huanyu Zhang +11
The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for…
cs.CV2025
Hierarchical Deep Fusion Framework for Multi-dimensional Facial Forgery Detection -- The 2024 Global Deepfake Image Detection Challenge
Kohou Wang, Huan Hu, Xiang Liu +4
The proliferation of sophisticated deepfake technology poses significant challenges to digital security and authenticity. Detecting these forgeries, especially across a wide spectr…
cs.CV2025★ 3 cited
A Large Vision-Language Model based Environment Perception System for Visually Impaired People
Zezhou Chen, Zhaoxiang Liu, Kai Wang +2
It is a challenging task for visually impaired people to perceive their surrounding environment due to the complexity of the natural scenes. Their personal and social activities ar…