Showing 2025Show all
2 papers · 1 filter
cs.CL2025
SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment
Guoxin Zang, Xue Li, Donglin Di +4
While Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle in industrial anomaly detection and reasoning, particularly in de…
cs.CL2025
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
Shibo Sun, Xue Li, Donglin Di +6
While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfac…