14 papers
Toward Generalizable Forgery Detection and Reasoning
Yueying Gao, Dongliang Chang, Bingyao Yu +5
Accurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models…
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
Zhonghao Yan, Muxi Diao, Yuxuan Yang +7
Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual p…
Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers
Shuo Zhang, Wenzhuo Wu, Huayu Zhang +8
Recent advances in diffusion models have significantly improved image editing. However, challenges persist in handling geometric transformations, such as translation, rotation, and…
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
Muxi Diao, Lele Yang, Hongbo Yin +5
Effective autonomous driving hinges on robust reasoning across perception, prediction, planning, and behavior. However, conventional end-to-end models fail to generalize in complex…
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
Muxi Diao, Yutao Mou, Keqing He +6
The safety of Large Language Models (LLMs) is crucial for the development of trustworthy AI applications. Existing red teaming methods often rely on seed instructions, which limits…
RotatedMVPS: Multi-view Photometric Stereo with Rotated Natural Light
Songyun Yang, Yufei Han, Jilong Zhang +4
Multiview photometric stereo (MVPS) seeks to recover high-fidelity surface shapes and reflectances from images captured under varying views and illuminations. However, existing MVP…