3 papers
cs.AI2025
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
Hongyi Jing, Jiafu Chen, Chen Rao +9
The existing Multimodal Large Language Models (MLLMs) for GUI perception have made great progress. However, the following challenges still exist in prior methods: 1) They model dis…
cs.CV2024
Attack Deterministic Conditional Image Generative Models for Diverse and Controllable Generation
Tianyi Chu, Wei Xing, Jiafu Chen +5
Existing generative adversarial network (GAN) based conditional image generative models typically produce fixed output for the same conditional input, which is unreasonable for hig…
cs.CV2024
PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style Mapping
Jiafu Chen, Wei Xing, Jiakai Sun +7
3D scene stylization refers to transform the appearance of a 3D scene to match a given style image, ensuring that images rendered from different viewpoints exhibit the same style a…