2 papers
cs.CV2026
HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes
Yujia Li, Yiqun Zhang, Zihan Cheng +7
Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfulness while misidentifying the attacked tar…
cs.CL2026
ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement
Mary Lynn Martin, Yifei Zhang, Martha Palmer +1
We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and d…