5 papers
DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
Xinrun Xu, Pi Bu, Ye Wang +7
Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in co…
Catastrophic Forgetting Mitigation via Discrepancy-Weighted Experience Replay
Xinrun Xu, Jianwen Yang, Qiuhong Zhang +3
Continually adapting edge models in cloud-edge collaborative object detection for traffic monitoring suffers from catastrophic forgetting, where models lose previously learned know…
Vulnerability of Text-to-Image Models to Prompt Template Stealing: A Differential Evolution Approach
Yurong Wu, Fangwen Mu, Qiuhong Zhang +8
Prompt trading has emerged as a significant intellectual property concern in recent years, where vendors entice users by showcasing sample images before selling prompt templates th…
High-Quality Pseudo-Label Generation Based on Visual Prompt Assisted Cloud Model Update
Xinrun Xu, Qiuhong Zhang, Jianwen Yang +4
Generating high-quality pseudo-labels on the cloud is crucial for cloud-edge object detection, especially in dynamic traffic monitoring where data distributions evolve. Existing me…
A Clustering Method with Graph Maximum Decoding Information
Xinrun Xu, Manying Lv, Zhanbiao Lian +4
The clustering method based on graph models has garnered increased attention for its widespread applicability across various knowledge domains. Its adaptability to integrate seamle…