6 papers
AnyStyle: A Single LoRA is Sufficient for Image-Guided Style Transfer
Yongwen Lai, Chaoqun Wang
Image-guided style transfer aims to apply the artistic characteristics of a style image to a content image while preserving its semantic structure and layout. Despite advances in d…
GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks
Wei Zhang, Chaoqun Wang, Zixuan Guan +5
Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Existing slice discovery approaches l…
Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing
Zhuohan Ouyang, Zhe Qian, Wenhuo Cui +1
Instruction-guided image editing requires balancing target modification with non-target preservation. Recently, flow-based models have emerged as a strong and increasingly adopted…
FusionEdit: Semantic Fusion and Attention Modulation for Training-Free Image Editing
Yongwen Lai, Chaoqun Wang, Shaobo Min
Text-guided image editing aims to modify specific regions according to the target prompt while preserving the identity of the source image. Recent methods exploit explicit binary m…
Unlock the Power of Unlabeled Data in Language Driving Model
Chaoqun Wang, Jie Yang, Xiaobin Hong +1
Recent Vision-based Large Language Models~(VisionLLMs) for autonomous driving have seen rapid advancements. However, such promotion is extremely dependent on large-scale high-quali…
Semantic-Supervised Spatial-Temporal Fusion for LiDAR-based 3D Object Detection
Chaoqun Wang, Xiaobin Hong, Wenzhong Li +1
LiDAR-based 3D object detection presents significant challenges due to the inherent sparsity of LiDAR points. A common solution involves long-term temporal LiDAR data to densify th…