6 papers
ARENA: Automated Red-Teaming for Large Audio Language Models
Jiaming He, Zhicong Huang, Tian Jin +5
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that…
A Smart Classroom Behavior Analysis Framework with a New Highly Congested Classroom Dataset
Wei Xu, Maoxiang Chu, Yuelong Fan +5
Student behavior detection is important for intelligent classroom analysis but remains challenging in large-class scenarios due to dense instance co-occurrence, asymmetric occlusio…
ChronoLock: Protecting Videos from Unauthorized Text-to-Video Personalization
Jiaming He, Jiashu Zhang, Guanyu Hou +4
Text-to-video (T2V) diffusion models have made it increasingly easy to synthesize realistic and temporally coherent videos, while recent personalization techniques allow such model…
HD-DinoMoE: A Class-Aware Hierarchical Dual Mixture-of-Experts Network for Scleral Anomaly Segmentation in Complex Acquisition Scenarios
Yinxiang Yu, Maoxiang Chu, Qi Niu +7
Traditional Chinese Medicine (TCM) ocular inspection provides empirical cues for assessing scleral surface anomalies, but its clinical use remains subjective and difficult to quant…
UTOPIA: Unlearnable Tabular Data via Decoupled Shortcut Embedding
Jiaming He, Fuming Luo, Hongwei Li +5
Unlearnable examples (UE) have emerged as a practical mechanism to prevent unauthorized model training on private vision data, while extending this protection to tabular data is no…
TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models
Jiaming He, Guanyu Hou, Hongwei Li +6
Text-to-Video (T2V) models are capable of synthesizing high-quality, temporally coherent dynamic video content, but the diverse generation also inherently introduces critical safet…