5 papers · 1 filter
Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs
Mingyu Yu, Lana Liu, Zhehao Zhao +2
The rapid advancement of Multimodal Large Language Models (MLLMs) has introduced complex security challenges, particularly at the intersection of textual and visual safety. While e…
Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method
Chao Huang, Benfeng Wang, Wei Wang +5
Recent progress in reasoning capabilities of Multimodal Large Language Models(MLLMs) has highlighted their potential for performing complex video understanding tasks. However, in t…
From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
Yuqing Shao, Yuchen Yang, Rui Yu +5
End-to-end multi-object tracking (MOT) methods have recently achieved remarkable progress by unifying detection and association within a single framework. Despite their strong dete…
UltraAD: Fine-Grained Ultrasound Anomaly Classification via Few-Shot CLIP Adaptation
Yue Zhou, Yuan Bi, Wenjuan Tong +3
Precise anomaly detection in medical images is critical for clinical decision-making. While recent unsupervised or semi-supervised anomaly detection methods trained on large-scale…
Measure Anything: Real-time, Multi-stage Vision-based Dimensional Measurement using Segment Anything
Yongkyu Lee, Shivam Kumar Panda, Wei Wang +1
We present Measure Anything, a comprehensive vision-based framework for dimensional measurement of objects with circular cross-sections, leveraging the Segment Anything Model (SAM)…