3 papers
cs.CV2025
MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operations
Jiang Wu, Sichao Wu, Yinsong Ma +4
Industrial accidents, particularly in high-risk domains such as surface and underground mining, are frequently caused by unsafe worker behaviors. Traditional manual inspection rema…
cs.CV2025
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
Meng Luo, Shengqiong Wu, Liqiang Jing +12
Recent advancements in large video models (LVMs) have significantly enhance video understanding. However, these models continue to suffer from hallucinations, producing content tha…
eess.IV2025
DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis
Minxi Ouyang, Lianghui Zhu, Yaqing Bao +10
Multimodal large models have shown great potential in automating pathology image analysis. However, current multimodal models for gastrointestinal pathology are constrained by both…