11 papers
When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy
Xiaofeng Tan, Jun Liu, Bin-Bin Gao +5
RLHF is widely used to align flow-matching text-to-image models with human preferences, but often leads to severe diversity collapse after fine-tuning. In RL, diversity is often as…
AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation
Xi Jiang, Yinjie Zhao, Zesheng Yang +1
Visual anomaly detection (VAD) is crucial in many real-world fields, such as industrial inspection, medical imaging, infrastructure monitoring, and remote sensing. However, the spe…
STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning
Zinuo Li, Yongxin Guo, Jun Liu +7
Human understanding of video dynamics relies on forming structured representations of entities, actions, and temporal relations before engaging in abstract reasoning. In contrast,…
AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison
Xi Jiang, Yue Guo, Jian Li +7
Multimodal Large Language Models (MLLMs) have achieved impressive success in natural visual understanding, yet they consistently underperform in industrial anomaly detection (IAD).…
Mastering Negation: Boosting Grounding Models via Grouped Opposition-Based Learning
Zesheng Yang, Xi Jiang, Bingzhang Hu +4
Current vision-language detection and grounding models predominantly focus on prompts with positive semantics and often struggle to accurately interpret and ground complex expressi…
ConsistentRFT: Reducing Visual Hallucinations in Flow-based Reinforcement Fine-Tuning
Xiaofeng Tan, Jun Liu, Yuanting Fan +7
Reinforcement Fine-Tuning (RFT) on flow-based models is crucial for preference alignment. However, they often introduce visual hallucinations like over-optimized details and semant…