30 papers
PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation
Leyi Sheng, Han Sun, Zhen Sun +4
As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. More importantly, T2I jailbreak evaluation i…
Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
Xusheng Liang, Lihua Zhou, Nianxin Li +8
Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their application to medical imaging rem…
MAT-Cell: A Multi-Agent Tree-Structured Reasoning Framework for Batch-Level Single-Cell Annotation
Yehui Yang, Zelin Zang, Xienan Zheng +9
Automated single-cell annotation is difficult when the most abundant genes are not the most discriminative ones, or when a target state is poorly covered by a fixed reference atlas…
UniDA3D: A Unified Domain-Adaptive Framework for Multi-View 3D Object Detection
Hongjing Wu, Cheng Chi, Jinlin Wu +3
Camera-only 3D object detection is critical for autonomous driving, offering a cost-effective alternative to LiDAR based methods. In particular, multi-view 3D object detection has…
Learning to Think Fast and Slow for Visual Language Models
Chenyu Lin, Cheng Chi, Jinlin Wu +2
When faced with complex problems, we tend to engage in slower, more deliberate thinking. In contrast, for simple questions we give quick, intuitive responses. This dual-system thin…
MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models
Siqi Ma, Jiajie Huang, Fan Zhang +5
Answering complex medical questions requires not only domain expertise and patient-specific information, but also structured and multi-perspective reasoning. Existing multi-agent a…